推荐理由:报道称苹果在 Siri 改造中采用定制 1.2T 参数 Google 模型,读者可了解其规模与端侧推理之间的取舍。
端侧 AI
全部主题跑在手机、电脑与边缘设备上的 AI:小模型、本地推理与端侧芯片的进展。
最新精选
第 21–28 条 · 共 28 条@kimmonismus@kimmonismus精选AI 评分7070 
@berryxia@berryxia精选AI 评分6868 引用dadabots (@dadabots)@dadabots🥳 Announcing Stable Audio 3 🍕 🏆 fastest music models ever 💻 runs on MacBookPro M-series 🧪 break it plz 🧠 LoRA finetune in < 1h 📷 Sm = faster, Medium = qualityer ⚡ 59x realtime on M5 Pro One-liner fast install: curl -LsSf dadabots.com/_/sa3-mac | bash Video
推荐理由:Stable Audio 3 给出在 Mac 本地运行音乐生成的安装方式与性能数据,可据此判断本地音乐生成工作流的可行性。
@vista8@vista8精选AI 评分6666 引用OpenBMB (@OpenBMB)@OpenBMB1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀 High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency. 🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget. ⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images. 🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090. Try the model here: 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM-V 🔭 Modelscope: modelscope.cn/models/OpenBMB… 🌐 Web Demo: huggingface.co/spaces/openbm… 📱 App Demo: github.com/OpenBMB/MiniCPM-V… Video
推荐理由:官方给出 1.3B 小模型处理高分辨率图像的编码成本与吞吐数据,可供端侧多模态选型参考。
@berryxia@berryxia精选AI 评分7272
引用Daniel Han (@danielhanchen)@danielhanchenWe released experimental MTP Qwen3.6 Unsloth GGUFs! Qwen3.6 27B MTP now runs at 140 tokens/s. Qwen3.6 35B-A3B MTP gets 220 tokens/s generation on a single GPU. Qwen3.6 27B and 35B-A3B have >1.4x speed-up over the original GGUFs without any change in accuracy. Guide + GGUFs + Benchmarks: unsloth.ai/docs/models/qwen3… In terms of average speedup, we see a 1.4x for dense models at draft tokens = 2 and for the MoE around 1.15 to 1.2x. We do not recommend more than 2 draft tokens because the acceptance rate drops precipitously from 83% to 50% with 4 draft tokens, and the forward passes for MTP become less beneficial. Use `--spec-type mtp --spec-draft-n-max 2` Thanks to Aman for github.com/ggml-org/llama.cp…!
推荐理由:原文给出单 GPU 实测速度、加速比与 draft tokens 甜点,可据此判断本地 30B 级模型的部署空间。
微信公众号(Mp2RSS 合集)精选AI 评分7777 MiniCPM-o 4.5 技术报告发布:全双工全模态 API 开放,RTX5070 即可实时运行
面壁智能联合 OpenBMB、清华大学 THUNLP 与 THUMAI 实验室发布 MiniCPM-o 4.5 技术报告,首次公开全双工全模态框架 Omni-Flow,并同步开放全模态全双工 API、在线 Demo、端侧安装包 Comni 与 Demo 仓库。
推荐理由:技术报告首次公开 Omni-Flow 的毫秒级时间轴对齐机制,读者可据此理解端到端全双工全模态的实现路径。
@TencentHunyuan@tencenthunyuan精选AI 评分6666 


推荐理由:1.8B 翻译模型量化到 1.25bit 后仅 440MB,可在手机离线运行,为端侧模型压缩提供参考。
Hugging Face Blog精选AI 评分6767 如何在 Chrome 扩展中用 Transformers.js 运行本地 Gemma 4 模型
Hugging Face 发布指南,讲解如何在 Chrome 扩展(Manifest V3)中用 Transformers.js 运行本地模型,并复现其 Gemma 4 浏览器助手的核心架构。
推荐理由:用一个已上线的浏览器扩展拆解 MV3 下的分层架构,可供搭建本地模型扩展时参考分工与数据边界。
Mistral AI精选AI 评分7777 Mistral 发布 Mistral 3 系列模型,含 Mistral Large 3 与 Ministral 3,均以 Apache 2.0 开源
Mistral 发布新一代模型系列 Mistral 3,包括 14B/8B/3B 的 Ministral 3 小模型和 sparse MoE 架构的 Mistral Large 3(41B 激活、675B 总参数),全部以 Apache 2.0 许可开源,提供 base、instruct 和 reasoning 变体。
推荐理由:官方公布完整模型规格、开源许可和部署路径,读者可据此评估在自建或端侧场景的可用性。