AI 导读
MiniMax 在 M3 博客中披露,模型约 24 小时内自主完成 CUDA FP8 GEMM 优化,期间提交 147 次基准测试、执行 1,959 次工具调用,把 Hopper FP8 峰值利用率从 7.6% 提升到 71.3%,实现 9.4 倍加速。
正文
CUDA kernel optimization in M3 Blog:
- FP8 GEMM: most compute-heavy and hardest-to-optimize part of inference; ~1–2 weeks for a senior team on Hopper.
- Cold start: only a task description, benchmark script, and non-runnable Triton skeleton—no reference implementation.
- ~24 hours autonomous: 147 benchmark submissions, 1,959 tool calls.
- Full path solo: baseline → autotune → bottleneck diagnosis → CUDA Graph → persistent kernel → host-side scheduling.
- Result: 7.6% → 71.3% Hopper FP8 peak utilization, 9.4× speedup.
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days在 X 查看被引用的帖子
来源:@MikaStars39 · x.com