AI 导读
又一个长时程基准测试,又一次提醒我们:当前 AI 可以工作数小时,但仍远未掌握任务。 EdgeBench 中最强的智能体在允许工作 12 小时后得分 51.3/100。 该基准为智能体提供持久可执行环境和反馈,涵盖 134 项任务。它们可以调试代码、运行模拟、检查验证结果、修改证明,并反复向隐藏评审提交产物。 也就是说,它们具备了我们通常认为智能体需要随时间改进的许多要素。 但它们仍难以将所有这些交互转化为持续进步。 该研究覆盖 6 个任务族、约 38,000 小时的交互。 当跨任务取平均性能时,改进遵循对数-S 型曲线,在全部 5 个模型上 R² ≥ 0.997:早期进展缓慢,随后是更陡的学习阶段,然后饱和。
正文
https://t.co/QBlkWurCo7
Another long-horizon benchmark, another reminder that current AI can work for hours and still remain far from mastering the task. The strongest agent in EdgeBench scored 51.3/100 after being allowed to work for 12 hours. The benchmark gives agents persistent executable environments and feedback across 134 tasks. They can debug code, run simulations, inspect validation results, revise proofs, and repeatedly submit artifacts to hidden judges. i.e. they had many of the ingredients we normally say agents need to improve over time. They still struggle to convert all of that interaction into sustained progress. The study covers roughly 38,000 hours of interaction across 6 task families. When performance was averaged across tasks, improvement followed a log-sigmoid curve with R² ≥ 0.997 across all 5 models: slow early progress, a steeper learning phase, then saturation.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com