微软等提出 ScholarEvolve,依据已发表的智能体研究而非自身失败日志来改进 harness,并将其拆分为工具使用、记忆管理、任务执行三个模块,通过主题建模从近期论文中提取各模块的改进策略再组合测试,且可持续加入新论文。
New paper from Microsoft and colleagues on evolving agent harnesses.
It's a really cool idea to evolve a harness from published research. Something I have also been testing for the past couple of months.
ScholarEvolve proposes harness changes based on published agent research rather than the agent's own failure logs.
It splits the harness into modules for tool use, memory management, and task execution.
It runs topic modeling over recent papers to identify distinct improvement strategies for each module, then implements and tests combinations.
New papers can be added over time.
With the model held fixed, Qwen3.5-27B goal completion on AppWorld Challenge rises from 49.6% to 63.6%, and GPT-5.4-mini on Tau2-Bench Telecom rises from 72.7% to 81.9%.
Paper: https://arxiv.org/abs/2609.40169
Chat with Paper: https://academy.dair.ai/papers/learning-from-research-toward-lifelong-agent-harness-evolution-2609.40169
来源:elvis · x.com