跳到正文
原文
Google DeepMind·· 2026-06-11精选AI 评分72

Google DeepMind 发布开源实验模型 DiffusionGemma,文本生成速度最高提升 4 倍

DiffusionGemma: 4x faster text generation

AI 导读

Google DeepMind 发布 Apache 2.0 开源的实验性文本扩散模型 DiffusionGemma,为 26B MoE 架构、推理时激活 3.8B 参数,在专用 GPU 上文本生成最高提速 4 倍,单张 NVIDIA H100 超过 1000 tokens/秒,量化后可放入 18GB 显存。

推荐理由

官方给出了具体吞吐数字、显存占用和适用场景边界,读者可以据此判断扩散文本生成适合哪些本地工作流。

来源:Google DeepMind · deepmind.google