Apple 论文《How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
论文用同一模型和硬件对比多种 harness,给出单智能体优于多智能体编排的具体数据,对构建 MLE agent 有直接参考。
New Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them.
It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session.
Those systems add search trees, memory layers, and specialist agent teams.
They were designed when a model could only write code, not run it.
This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered.
With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%.
– arxiv. org/abs/2609.40303
Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"
来源:Rohan Paul · x.com