跳到正文
@EpochAIResearch· @EpochAIResearch · X·· 2026-08-27AI 评分46
AI 导读

根据 @ArtificialAnlys 对 Gemma 4 31B 的基准测试,Nvidia 的 Groq 3 LPX 系统在小模型快速解码方面表现出色(>3,400 tok/s)。LPX 通过使用少量(128 GB)超高速 SRAM 替代 HBM 来实现这一点。 Nvidia 提议将 LPU 与 GPU 结合,把这一速度带到前沿规模的模型上。下面来看看这可能如何实现。

正文

Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to a @ArtificialAnlys benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM.

Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models. A look at how that (might) work below.

来源:@EpochAIResearch · x.com