跳到正文
@bleysg· @bleysg · X·· 2026-06-02AI 评分52
AI 导读

作者 @bleysg 在 MiniMax M3 上运行 DeepSWE,113 项中通过 15 项,计入 1.5 倍超时则为 19 项。这是针对 MiniMax 新发布开源权重模型 M3 的独立测试,完整报告见 entrpi.github.io/misc/deep-swe-minimax-m3/。MiniMax 官方称 M3 为其首个同时具备三项前沿能力的开源权重模型,SWE-Bench Pro 得分为 59.0%。

正文

Since everyone is asking, I ran DeepSWE on MiniMax M3.

Here is the lowdown. 15 of 113 passed!

19 if you count the 1.5x overtime I gave just to see.

Full report: entrpi.github.io/misc/deep-s…

引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days
在 X 查看被引用的帖子

来源:@bleysg · x.com