SemiAnalysis 称 OpenAI 新芯片 Jalapeño 将于 2026 年底前部署在其自有算力中,在约 100 tok/s/用户的交互性下可达约 11M tok/s/MW,约为表现最好的 Blackwell 级配置的 2 倍。
Semianalysis on OpenAI's new Jalapeño chips that will be deployed inside its own compute by the end of 2026.
"Jalapeño smokes every other chip. All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP"
Jalapeño shifts the entire latency–efficiency Pareto frontier upward: at ~100 tok/s/user it delivers ~11M tok/s/MW, roughly 2× the best Blackwell-class configs at comparable interactivity.
The kicker: that’s STP (Single-Token Prediction) with no speculative decoding, while the competing curves are already using MTP.
来源:@rohanpaul_ai · x.com