OpenAI 称其新款 Jalapeño 芯片在匹配 DeepSeek R1 解码速度时,每千瓦吞吐达到 NVIDIA GB300 的 104.3 倍。
MASSIVE: OpenAI just claimed its new Jalapeño chips delivered 104.3x more throughput per kilowatt than NVIDIA GB300 at matched DeepSeek R1 decoding speed.
The figure comes from GB300's previous-best time-between-token speed: at 169.41 tok/s/user, Jalapeño produced 12,258 mixed tokens/s/kW versus 118 for GB300.
The tests used SemiAnalysis's public InferenceX benchmark, which measures full-request throughput, power efficiency, and latency.
- Jalapeño is rated at 700W, compared with 1,200W for GB200 and 1,400W for GB300
- OpenAI attributes the efficiency to keeping KV-cache and model state local, reducing data movement as inference switches between compute-heavy prefill and memory-bound decoding.
- OpenAI plans deployment by year-end, with Gen 2 already well underway and Gen 3 taking shape
- AI also helped move Jalapeño from initial design to tapeout in nine months, and selected AI-generated GPT-OSS kernels ran 1.5 to 1.8x faster than human-expert implementations on selected blocks.
来源:@rohanpaul_ai · x.com