@vista8· @vista8 · X·· 2026-05-15精选AI 评分66
AI 导读
面壁智能发布 1.3B 参数的视觉模型 MiniCPM-V 4.6,面向消费级和移动硬件,已在 Hugging Face、GitHub 和 ModelScope 上线。该模型采用 LLaVA-UHD v4 技术,官方称将视觉编码成本降低 55%。官方还称其在多模态和 Artificial Analysis 基准上超过 Gemma4-E2B-it 和 Qwen3.5-0.8B,TTFT 为 75.7ms、比 Qwen3.5-0.8B 快 2.2 倍。原文作者表示在 Hugging Face 看到该模型论文,准备抽空测试。
推荐理由
官方给出 1.3B 小模型处理高分辨率图像的编码成本与吞吐数据,可供端侧多模态选型参考。
正文
前几天在Huggingface看到模型论文了。
面壁智能的MiniCPM-V 4.6 ,竟然只有1.3B的视觉模型。
看Benchmark效果有点强,抽空测试下。
1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀 High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency. 🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget. ⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images. 🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090. Try the model here: 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM-V 🔭 Modelscope: modelscope.cn/models/OpenBMB… 🌐 Web Demo: huggingface.co/spaces/openbm… 📱 App Demo: github.com/OpenBMB/MiniCPM-V… Video在 X 查看被引用的帖子
来源:@vista8 · x.com