跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 27 天前AI 评分42
AI 导读

TPUv7 内置专用硬件加速单元 SparseCore,负责 MoE 推理中的数据搬运,将各专家的 token 聚集成连续分组,TensorCore 则专注执行专家矩阵乘法,这一分工让吞吐量提升 12%。结合其他优化与 TPU 更低的 TCO,在 InferenceX 上 TPU 每美元性能最高比 Blackwell Ultra 高出 50%。

正文

TPUv7 has a specialized hardware-accelerated unit called the SparseCore. TPU's specialized SparseCore handles data movement, gathering each expert’s tokens into contiguous groups, while the TensorCore is left to run the expert matrix multiplications. When using SparseCore for rearranging expert inputs into the MoE kernel, it results in 12% better throughput.

Combined with other optimizations & TPU's lower TCO, TPU can achieve up to 50% better perf per dollar than Blackwell Ultra, as seen on InferenceX.

来源:@SemiAnalysis_ · x.com