跳到正文
原文
@googleaidevs· @googleaidevs · X·· 2026-06-06精选AI 评分71
AI 导读

Google 发布 Gemma 4 QAT 检查点,可在消费级 GPU 和移动设备上本地运行且质量损失很小。其中 GGUF(Q4_0)检查点覆盖各尺寸与 drafter 模型以追求本地性能,定制移动方案通过混合精度把 Gemma 4 压缩到 1GB 以内,包含针对性 2-bit 解码层、优化 KV cache 和静态激活。

推荐理由

文章列出 Gemma 4 QAT 的 GGUF 与移动端两种量化方案,可用于判断本地部署的体积与质量取舍。

正文

New @GoogleGemma 4 QAT (Quantization-Aware Training) checkpoints are here, so you can run models locally on consumer GPUs and mobile devices with minimal quality loss.

What’s new:

🔹 GGUF (Q4_0): Checkpoints: Max local performance across all sizes and drafter models

🔹 Custom Mobile Schema: We shrunk Gemma 4 down to less than 1GB for mobile devices by using a custom mixed precision schema designed for edge hardware (featuring targeted 2-bit decoding layers, optimized KV caches, and static activations)

By simulating compression during training rather than after (Post-Training Quantization), we've drastically reduced the memory footprint and accelerated decode speeds while preserving reasoning quality. blog.google/innovation-and-a…

来源:@googleaidevs · x.com