跳到正文
@kimmonismus· @kimmonismus · X·· 2026-06-06AI 评分60
AI 导读

Google DeepMind 发布 Gemma 4 QAT 系列模型,采用量化感知训练,在压缩内存需求的同时比常规训练后量化保留更多质量。该系列支持 Q4_0 格式和一种新的移动端专用量化格式,Gemma 4 E2B 约 1GB 内存即可运行,纯文本版本所需内存甚至低于 1GB,让手机、笔记本、边缘设备和消费级 GPU 上的本地 AI 更实用。

正文

Google DeepMind released new Gemma 4 QAT models that make the model family much more efficient for local, on-device use.

Using Quantization-Aware Training, the models are trained with compression in mind, which reduces memory needs while preserving more quality than standard post-training quantization. The release includes support for the popular Q4_0 format and a new mobile-specialized quantization format.

Gemma 4 E2B can now run with around 1GB of memory (!), and the text-only version can even require less than 1GB (!). That makes local AI on phones, laptops, edge devices, and consumer GPUs far more practical.

Really cool to see.

来源:@kimmonismus · x.com