跳到正文
原文
@kimmonismus· @kimmonismus · X·· 2026-08-26精选AI 评分71
AI 导读

OpenAI 在关于 Jalapeño 芯片的博客中披露,GPT-Astra 用 Codex 编写并优化底层 kernel,两个月内把三个原本不在计划内的开源权重模型带到该芯片上的高性能。在选定的 attention 和 MoE 模块上,AI 生成的实现比现有人工专家代码快 1.5–1.8 倍,这些数字只适用于选定模块而非完整模型。引用的推文还提到,Jalapeño 在 GPT-OSS 120B、DeepSeek R1 670B 和 Kimi K2.5 1T 上实现 1.5–1.9 倍每瓦 AI 吞吐与 1.7–3.6 倍更低的端到端延迟,OpenAI 计划 2026 年底开始部署。

推荐理由

原文给出 GPT-Astra 用 Codex 生成 kernel 的具体加速数据,读者可据此了解 AI 编写底层算子的表现。

正文

interesting detail in OpenAIs blogpost to "Jalapeño":

GPT-Astra used Codex to write and optimize low-level kernels that brought three previously unplanned open-weight models to high performance on OpenAI’s Jalapeño chip within two months.

For selected attention and mixture-of-experts blocks, its implementations ran 1.5–1.8× faster than existing human-expert-written code.

looks like astra is a really really good model.

引用@kimmonismus@kimmonismus
Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing! For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance. The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads. OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape. Probably thats why Tibo said that in 1-2 years 750token/s will be the default
在 X 查看被引用的帖子

来源:@kimmonismus · x.com