跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-24AI 评分48
AI 导读

一项多智能体代码审查新研究提出 Adversarial Review:由主编码智能体编写、审查者评估、批评者审计审查,三个智能体在编辑前完成结构化冲突。该方法在 LiveCodeBench 上以三个智能体击败五智能体基线;在 SWE-PRBench 上,朴素版本因智能体在证据不足时趋于一致而暴露失败模式,将分歧设为明确指令后取得测试方法中最高的 F1。

正文

Great paper on multi-agent systems for code review.

It's challenging to know how many coding agents to use to address a problem.

The default fix for weak agentic code review is more agents. In turns out that scaling agents to a large number gives diminishing returns on repository-level tasks.

This new work tries structured conflict instead. Adversarial Review runs three agents. A main coding agent writes, a reviewer evaluates, and a critic audits the review before any edit are done.

On LiveCodeBench it beats a five-agent baseline while using three agents.

On SWE-PRBench the naive version exposed a failure mode. The agents converged on agreement without enough evidence behind it. Making disagreement an explicit instruction recovered the highest F1 among tested methods.

They also find that cooperative review works when the disagreement is minimal, structured, and grounded in evidence.

Paper: https://t.co/NY1gcqajI0

Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

来源:@omarsar0 · x.com