跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-29AI 评分43
AI 导读

又一篇论文清晰地揭示了AI的长时程问题。 Long-Horizon-Terminal-Bench,一旦任务延伸到数百步,即便是前沿智能体也举步维艰。 在46个长终端任务上测试17个前沿模型,平均通过率仅6.4%。 在严格的全完成度评分下,17个模型中有10个任务解决数为零。失败运行中,79%以时钟到期、智能体仍在工作告终。 即便最好的模型也只有28.3%,每10个任务中有7个未完成。 这些模型能执行大量局部合理的步骤,却无法可靠地将数百个步骤转化为一个完成的结果。

正文

https://t.co/Nvj8u0SF91

引用@rohanpaul_ai@rohanpaul_ai
Another paper that so clearly exposes AI’s long-horizon problem. Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps. Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%. And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working. Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished. The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com