AI 导读
NVIDIA LPU 支持 3 种解耦推理: 1. Rubin Prefill + LPU Decode,实现最快交互性 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN,适用于曲线中段 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter,适用于曲线中左段 对于低交互性场景,纯 Rubin 仍然胜出。期待看到 Rubin + LPU 在 AgentX 等开源智能体基准上的性能曲线。
正文
NVIDIA LPU supports 3 types of disaggregated inferencing:
1. Rubin Prefill + LPU Decode for the fastest interactivity
2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve
3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve
For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
来源:@SemiAnalysis_ · x.com