Google 发布 Gemma 4 12B,一款不含多模态编码器的统一模型,可直接在笔记本上运行。该模型定位在移动端 E4B 与更大的 26B MoE 模型之间,视觉和音频输入直接进入 LLM 主干,原生支持音频,采用 Apache 2.0 许可。官方称在 16GB VRAM 下可本地运行复杂多步工作流,性能接近 26B 模型。
官方给出无编码器架构与 16GB VRAM 本地运行的定位,读者可据此判断端侧多模态模型的选型边界。
We’re launching Gemma 4 12B: Our unified, encoder-free model that brings powerful multimodal intelligence straight to your laptop 🚀
The model bridges the gap between our mobile E4B model and larger 26B MoE models, packaging frontier-class reasoning and native audio into a highly optimized footprint, all under a permissive Apache 2.0 license.
Here’s what makes it unique:
+ Encoder-Less Architecture: We removed the multimodal encoders. The vision and audio inputs flow directly into the LLM backbone.
+ Agentic Performance (16GB VRAM): Run complex, multi-step workflows locally, with performance nearing our 26B model.
来源:@googleaidevs · x.com