引用@kimmonismus@kimmonismus
Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency.
Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing!
For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance.
The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads.
OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape.
Probably thats why Tibo said that in 1-2 years 750token/s will be the default