Google 发布 Gemma 4 12B 开源模型,采用 Apache 2.0 许可,可在 16GB 显存笔记本上本地运行,支持智能体推理、视觉与音频,作者称其质量接近 Google 的 26B 模型。
它把视觉与音频编码器并入主干,让 12B 模型能在 16GB 显存本地运行,读者可据此判断端侧多模态的门槛变化。
Gemma 4 12B shipped today under the label "encoder-free."
A local 12b model that shows really good results. I'm a big fan of Gemma Gemma 4 12B is out: a dense, fully open model (Apache 2.0) that runs on a 16GB laptop and does agentic reasoning, vision and audio at a quality Google puts near its 26B model.
The reason a 12B can pull this off: Google removed the separate vision and audio encoders and feeds both straight into the model, which keeps the memory footprint small enough for consumer GPUs.
For on-device assistants and private coding agents, that lowers the bar a lot. always look forward to the updates. 12b is a good sweet spot in terms of size.
a few facts:
Vision: the 550M encoder (27 transformer layers) is now a 35M embedder, one matmul on 48x48 pixel patches. Roughly 15x smaller.
Audio: the 300M encoder (12 conformer layers) is gone. Raw 16kHz audio cut into 40ms frames, projected straight into the LLM. So encoding didn't vanish, it collapsed into the backbone.
The payoff is real: one shared set of weights, so you LoRA-tune vision, audio and text in a single pass.
Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run locally with just 16GB of VRAM. It’s open and accessible for everyone to use under a permissive Apache 2.0 license. This is all made possible by our new, unified architecture that removes separate multimodal encoders. Here’s how we did it 🧵在 X 查看被引用的帖子
来源:@kimmonismus · x.com