跳到正文
Rohan Paul· @rohanpaul_ai · X·· 4 小时前AI 评分63
AI 导读

Meta 论文 RankEvolve 发现,让两个不同编码智能体互相审查补丁,比给单一智能体更大预算更能捕捉静默缺陷。在同一任务上混合使用 Claude Code 和 Codex,在同等花费下将完全正确的补丁比例从 45.8% 提升到 62.5%。

正文

New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget.

Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.

Agents editing real training code can leak test data, break a gradient, or miswire a flag.

The code still runs, so you burn GPU hours and get numbers that look valid.

The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase.

– arxiv. org/abs/2609.39551

Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"

来源:Rohan Paul · x.com