跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 2026-08-28AI 评分42
AI 导读

更大的故事在于这解锁了什么:异构解耦推理。晶圆是一台解码机器,其 roofline 对计算密集型的 prefill 表现不佳。有了新的 I/O 模块,Cerebras 可以在 prefill-decode 和 attention-FFN 解耦架构中与基于 HBM 的 XPU(已公布的合作伙伴是 AMD 和 Trainium)配对。这是绕过 44GB SRAM 上限的路径,让 HBM 系统承载晶圆无法容纳的内容。 完整文章👇️ (3/3) https://t.co/UV2hYz6d7j

正文

The bigger story is what this unlocks: heterogeneous disaggregated inference. The wafer is a decode machine, and its rooflines are poor for compute-bound prefill. With the new I/O module, Cerebras can pair with HBM-based XPUs (AMD and Trainium are the announced partners) in both prefill-decode and attention-FFN disaggregated setups. That's the path around the 44GB SRAM ceiling, letting HBM systems hold what the wafer can't.
Full Article👇️ (3/3) https://t.co/UV2hYz6d7j

来源:@SemiAnalysis_ · x.com