微软提出智能体 harness 自动优化方法 ActiveSaddler,不再喂固定任务列表,而是追踪失败模式、优先修复最值得解决的失败,或尝试未见任务以发现新失败。同一优化器下,GAIA2 测试通过率提升 4.4 个百分点,Terminal-Bench 2.0 提升 7.5 个百分点;达到 58.5% GAIA2 dev 准确率仅花 $298,固定顺序则需 $1,360。
New Microsoft paper on Automated harness optimization for agents.
Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets.
But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list.
ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones.
On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order.
If you auto-tune an agent, aim your run budget at the failures that are still open.
来源:Rohan Paul · x.com