EdgeBench 长时程基准测试显示,最强智能体在允许工作 12 小时后仅得 51.3/100 分。该基准为智能体提供持久可执行环境和反馈,覆盖 134 项任务,可调试代码、运行模拟、检查验证结果、修改证明并反复向隐藏评审提交产物。
Another long-horizon benchmark, another reminder that current AI can work for hours and still remain far from mastering the task.
The strongest agent in EdgeBench scored 51.3/100 after being allowed to work for 12 hours.
The benchmark gives agents persistent executable environments and feedback across 134 tasks. They can debug code, run simulations, inspect validation results, revise proofs, and repeatedly submit artifacts to hidden judges.
i.e. they had many of the ingredients we normally say agents need to improve over time.
They still struggle to convert all of that interaction into sustained progress.
The study covers roughly 38,000 hours of interaction across 6 task families.
When performance was averaged across tasks, improvement followed a log-sigmoid curve with R² ≥ 0.997 across all 5 models: slow early progress, a steeper learning phase, then saturation.
"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com