跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 29 天前AI 评分61
AI 导读

微软 AI 的新论文研究了 LLM 智能体在长程多步任务中的退化,跨 9 个模型的测试显示成功率通常随依赖步数增加而下降。在 ToolQA 上,短程近乎完美的模型到 16 步时成功率降到 0-33%;上下文长度并非主因,缩短上下文反而让退化更严重。论文建议按实际工作流长度测试智能体、衡量单步可靠性,并在坏步骤影响后续之前加入检查或检查点。

正文

New Microsoft paper. Long agent runs expose failures that short benchmarks miss. Agents can look reliable at 2 or 4 steps and fall apart by 16.

every agent step has some chance of going wrong, and those small errors compound as the workflow gets longer.

Across 9 models, success usually dropped as the number of dependent steps increased.

On ToolQA, models that were near-perfect on short runs fell to just 0-33% success by 16 steps.

Long context was not the main driver: shortening the context made the decline worse, so blindly trimming history is not a reliability fix.

For builders, the recommendation is straightforward: stop treating a benchmark pass rate as proof that an agent is production-ready.

Test agents at the workflow lengths you actually expect, measure per-step reliability, and add checks or checkpoints before a bad step poisons everything that follows.

来源:@rohanpaul_ai · x.com