跳到正文
@omarsar0· @omarsar0 · X·· 14 天前AI 评分54
AI 导读

Google 的 RRSI 论文提出用正则化约束智能体 harness 的自动进化,该方法在 evolve split 上得 90.5 分,并在 JobBench、GDPval 和 APEX-Agents 三个分布外基准上提升 3.5 至 4.7 分。

正文

Must-read paper from Google on self-improving agent harnesses.

If you auto-optimize your agent's harness, your eval score can go up while the agent gets worse on real tasks.

This paper shows how to prevent that.

Of five harness-evolution methods compared on agentic workspace tasks, RRSI scored the lowest on the tasks it evolved against and highest on all three out-of-distribution benchmarks.

Automated harness evolution proposes edits to prompts, control flow, tools and memory, keeps the ones that raise the score, and repeats.

The authors show this overfits the training tasks. Meta-Harness reached 93.0 on the Harvey LAB evolve split but gained only 0.3 to 1.5 points on JobBench, GDPval and APEX-Agents.

RRSI adds regularization on both sides of the loop.

The proposer gets an edit budget that shrinks over time and is pushed toward directions it has not tried. A critic rejects benchmark-specific edits, and a pruner removes edits that are too small, too costly or no longer useful.

RRSI scored 90.5 on the evolve split and gained 3.5 to 4.7 points on the three held-out benchmarks. In the ablation, unregularized evolution used 3.80M tokens per trial against 2.42M for RRSI. With Gemini 3.5 Flash, RRSI raised Terminal-Bench 2.1 from 64.6 to 78.7 and carried a 2.2-point gain over to SWE-bench Verified.

Paper: https://t.co/SlfjDg96VI

Chat with Paper: https://t.co/gBotiH6Jfq

来源:@omarsar0 · x.com