NVIDIA 宣布 Groq 3 LPX 已进入量产,为 Vera Rubin 平台新增专用 token 生成加速器。NVIDIA 称在 Artificial Analysis 基准测试中,以 100,000-token 上下文运行 Gemma 4 31B 达到每秒 3,400 output tokens,是该模型已记录的最快结果。
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated token-generation accelerator to the Vera Rubin platform.
NVIDIA says it reached 3,400 output tokens per second running Gemma 4 31B with a 100,000-token context in Artificial Analysis benchmarking, the fastest recorded result for that model.
The company also claims 4x faster responsiveness than the nearest alternative platform for agents and latency-sensitive workloads.
Nebius will be the first AI cloud to deploy Groq 3 LPX through its Token Factory, followed by Groq itself.
Intelligence too fast to meter.
来源:@kimmonismus · x.com