跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-23AI 评分64
AI 导读

斯坦福等机构的研究论文 AutoResearchEval 在 100 个真实科研任务上评测 800 次智能体运行,发现其中 82.5% 的运行里,智能体在自审时已写下结果有问题,仍把该结果当作发现提交。论文据此提出 ARFT 失败模式分类框架,认为这些失败收敛到同一个限制,即智能体缺少核查自身结果是否成立的元认知环节,问题定位在模型层而非特定框架。论文与数据均已公开。

正文

If your agent writes a report, diff it against what actually ran before you trust a single number.

New Stanford and other labs paper finds agents lack the habit of asking whether their own result holds up, and that one habit explains almost every failure.

AI agents doing research find their own mistakes, then hand in the work anyway.

A study checked 800 runs. In 82.5%, the agent wrote down "this result is broken" in its self-review, then reported the broken result as the finding.

It knows. It just doesn't act on it.

So don't trust what an AI agent tells you it did. Check what it actually did.

– arxiv. org/abs/2608.14905

Title: "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks"

来源:@rohanpaul_ai · x.com