跳到正文
原文
elvis· @omarsar0 · X·· 3 小时前AI 评分53
AI 导读

Google 发布 VeriHarness 论文,把同一基座模型变成智能体验证器,用于长程任务结果校验。

正文

Banger paper from Google.

It's standard practice to sample several agent rollouts and trust the answers they agree on.

This Google paper shows that agreement can hide shared errors, while disagreement often points to the correct alternative.

VeriHarness turns the same base model into an agentic verifier with two jobs.

One resolves claims where rollouts disagree by checking workspace evidence.

The other challenges claims that every rollout agrees on and looks for requirements they all missed.

Across five long-horizon benchmarks, it gives the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.

The authors also release about 26,000 rollouts.

Paper: https://arxiv.org/abs/2610.00972

Chat with Paper: https://academy.dair.ai/papers/veriharness-scaling-agentic-verification-for-long-horizon-tasks-2610.00972

来源:elvis · x.com