Long-Horizon-Terminal-Bench 测试了 17 个前沿模型在 46 个长终端任务上的表现,平均通过率仅 6.4%。严格全完成评分下,17 个模型中有 10 个零通过;失败运行中 79% 因超时而中断。表现最好的模型也仅 28.3%,10 个任务中有 7 个未完成。
Another paper that so clearly exposes AI’s long-horizon problem.
Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps.
Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%.
And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working.
Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished.
The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.
"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com