MiniMax 回顾了与 Together AI 团队围绕 M3 的直播分享,重点介绍 MSA(MiniMax Sparse Attention)机制。MSA 不压缩 KV 缓存,而是保留真实未压缩的 KV 并做块级 top-K 选择,使 1M 上下文窗口可行,注意力内核的每解码墙钟耗时占比从上代的约 30% 降至约 5%。
We wrapped a live session on M3 yesterday with the @togethercompute team & our researchers @zpysky1125 and @HaohaiSun
A few highlights 🧵
1. MSA (MiniMax Sparse Attention) is the star ⭐️. Unlike CSA/HCA, which compress the KV cache, MSA keeps the real, uncompressed KV and does block-level selection with a small top-K. That's how the 1M context window stays tractable.
2. The efficiency win is huge. In our previous generation, ~30% of per-decode wall-clock time went to the attention kernel. With MSA that now drops to ~5%. Big gains for long-context generation.
3. M3 isn't just a coding model. Natively multimodal (image + video in), ability to handle long-horizon agentic tasks, and even operate a desktop computer. People are already throwing game-dev + Minecraft-style builds at it (Unity included) and it's holding its own.
4. M3 can self-evaluate on vision-coding tasks: it builds a website or SVG, browses and inspects its own rendered output, judges it, and iterates - grading work visually.
5. We're also seeing junior-analyst-level performance on finance tasks; something we haven't even showcased publicly yet.
6. What's next: harder long-horizon / multi-file tasks in future releases, scaling data + post-training (RL) compute toward pre-training scale, and going deeper into finance, legal & bio.
Thanks to everyone who joined 🙏
Try M3 link in the comments👇
MiniMax M3 is live and Together AI is powering its inference 🚀 Tomorrow at 6pm PT we're going live on X Spaces with the teams behind the model and the infrastructure to give you a deep dive. nitter.net/i/spaces/1nxeLLDDBEaJX在 X 查看被引用的帖子
来源:@MiniMax_AI · x.com