跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 2026-09-01AI 评分28
AI 导读

NVIDIA LPU 支持 3 种解耦推理: 1. Rubin Prefill + LPU Decode,实现最快交互性 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN,适用于曲线中段 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter,适用于曲线中左段 对于低交互性场景,纯 Rubin 仍然胜出。期待看到 Rubin + LPU 在 AgentX 等开源智能体基准上的性能曲线。

正文

NVIDIA LPU supports 3 types of disaggregated inferencing:

1. Rubin Prefill + LPU Decode for the fastest interactivity
2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve
3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve

For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.

来源:@SemiAnalysis_ · x.com