跳到正文
@MikaStars39· @mikastars39 · X·· 2026-06-01AI 评分60
AI 导读

MiniMax M3 自主运行近 12 小时,复现了 ICLR 2025 杰出论文奖作品《Learning Dynamics of LLM Finetuning》,期间自行产出 18 次提交和 23 张实验图。它复现出 SFT 阶段预测的概率趋势,观察到 DPO 实验核心的 squeezing 效应,并验证了原论文提出的 Extend 缓解方法。

正文

Key takeaway from the M3 blog: M3 independently reproduce an ICLR 2025 Outstanding Paper Award winner, "Learning Dynamics of LLM Finetuning."

M3 ran autonomously for nearly 12 hours, producing 18 commits and 23 experimental figures on its own, and got the core experiments working:

- it matched the predicted probability trends in the SFT stage

- clearly observed the squeezing effect central to the DPO experiments

- validated the Extend mitigation method proposed in the original paper.

引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days
在 X 查看被引用的帖子

来源:@MikaStars39 · x.com