跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分43
AI 导读

论文《Adversarial Review》提出冻结代码、由 reviewer 写审查、critic 审计的结构,只有定稿审查才回传修改。在 LiveCodeBench 上 3 智能体达 87%,高于 5 智能体的 82%;但在真实 PR 审查中仅 0.457 F1 垫底。要求 critic 说明异议是引用代码还是仅凭直觉、reviewer 必须用代码回应后,分数升至 0.533 并登顶。

正文

2 AI agents set to check each other's work will usually end up agreeing, whether or not the code is right.

So the thing to add is not another reviewer but a rule about what an objection has to contain before either side is allowed to drop it.

This paper shows a reviewer plus a critic beating much larger review teams at writing code, then failing at reviewing code until that rule is in place.

The structure is small: the code stays frozen while a reviewer writes a review and a critic audits it, and only the settled review goes back for edits.

On LiveCodeBench it reaches 87% with 3 agents, against 82% for a 5-agent version.

On real pull-request review it lands last, at 0.457 F1. One prompt change fixes that: the critic must state whether its objection cites code or is only a hunch, and the reviewer has to answer with code either way, which takes it to 0.533 and the top of the set.

So the second agent only helps when agreement has to be paid for with code evidence; without that rule it mostly ratifies the first.

– arxiv. org/abs/2608.18167

Title: "Adversarial Review: Structured Disagreement for Grounded Agentic Code Review"

来源:@rohanpaul_ai · x.com