跳到正文
@kimmonismus· @kimmonismus · X·· 2026-05-26AI 评分36
AI 导读

MiniMax 预告 M3 的稀疏注意力架构,在 1M tokens 下对比 M2 实现 9.7 倍预填充加速和 15.6 倍解码加速。该方案采用两阶段方法:先用轻量索引分支做块选择,再仅对相关 KV 块执行稀疏注意力。MiniMax 此前因高效注意力尚未达到生产可用,在 M2 上刻意回归了全注意力。

正文

MiniMax just teased their Sparse Attention architecture for M3. The benchmarks show 9.7x prefilling speedup and 15.6x decoding speedup at 1M tokens vs M2.

MiniMax deliberately went back to full attention for M2 because efficient attention wasn't production-ready. Their pretrain lead wrote a whole blog post about it in March. Now they're showing a new two-stage approach, lightweight index branch for block selection, then sparse attention only on relevant KV blocks.

Really interesting. And tbh I'm always happy when open source receives new wins.

引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI
#MSA #OpenSource #M3 🫣😎
在 X 查看被引用的帖子

来源:@kimmonismus · x.com