跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-30AI 评分38
AI 导读

在 WeaveBench 的 114 个 GUI-CLI 混合任务上,官方报告的最佳通过率仅为 41.2%,说明长程智能体可靠性并未随模型变强而到来。当前模型能完成局部步骤,但需要显式审计的任务状态才能保持全程可控,否则会因历史记录无法可靠区分已完成、失败和待办事项而失败。

正文

Long-horizon agent reliability has not arrived yet with better models.

On WeaveBench's 114 hybrid GUI-CLI tasks, the best officially reported pass rate is 41.2%.

Current models can handle local steps, but need explicit audited state to keep the full task on track.

Long-horizon agents can solve individual steps and still fail because history becomes unreliable about what is finished, what failed, and what remains.

LongHorizon-Harness treats this as a task-state problem.

this paper argues that reliability over that span depends as much on the harness around the model as on the model itself.

引用@rohanpaul_ai@rohanpaul_ai
"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com