Meta、Duke 与加州大学的研究者提出将 Agent harness 的自动搜索拆分为多个专门分支,每个分支保留自己更擅长的问题并记录有效经验,再用路由器按任务分配最合适的分支,在全部 4 个测试设置中均超过 Meta-Harness。
New paper from Meta, Duke, California Univ on self-Improving Agent's Harness Optimization
When an AI tunes your agent's harness, split the search into specialized branches and route each task to the best fit, which beat Meta-Harness in all 4 test settings.
Giving each tuning branch its own problems and its own notes on what worked produced harnesses with different strengths, and a router turned those strengths into higher scores.
A harness is the code around an LLM that controls its tools, retrieval and self-checks. Meta-Harness has an AI rewrite it in a loop, but every version is scored on the same problems, so search sticks to 1 path.
Each of 2 branches keeps the practice problems it solves better than the other and writes its own notes on what worked. In math, 1 branch learned to verify answers, while the other learned to build full derivations.
With Gemini 3 Flash on Olympiad math, accuracy rose from 46.0% with Meta-Harness to 62.0%.
If you auto-tune agents, keep several specialized harnesses and route between them rather than betting on 1 winner.
来源:Rohan Paul · x.com