Microsoft 与多所机构研究者提出 AutoSaddler,把 Agent harness 视为代码,从失败执行轨迹中离线学习补丁来替代手工调优。它分批运行任务、诊断失败原因,生成针对提示词、工具配置和控制逻辑的结构化补丁,只在通过验证后保留更新。
Impressive new paper from Microsoft and colleagues.
Harness design is still hand-tuned almost everywhere. This work present an automated loop to optimize the harness.
They introduce AutoSaddler, which treats the agent harness as code and learns to patch it offline from failure traces.
It runs mini batches of tasks, diagnoses what broke, generates structured patches to prompts, tool configurations, and control logic, then keeps an update only if it survives validation.
Gains of 9.0 points on GAIA2, 9.6 on SWE-Bench Pro, and 10.0 on Terminal-Bench 2.0 over the corresponding base harnesses.
Deep debugging beats shallow reflection, targeted edits beat unconstrained editing, and generalization-aware selection beats repairing the one trajectory in front of you.
Paper: https://t.co/PoDahO6rmz
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com