AI 导读
我认为“时间跨度”正在成为讨论智能体能力更有用的方式之一。 在 OpenAI,随着任务变长,无需人工干预的成功率急剧下降,从 15 分钟以内任务的 86% 降至最长的 64-128 小时区间的仅约 16%,而需要干预的成功运行变得更加常见。 一个模型可以在局部极其强大,但在长执行轨迹上仍然不可靠。 不是 token。不是基准分数。系统能在人类必须介入之前对任务保持有效控制多久? 也许我们正在走向一种基准,类似 每个完成任务所消耗的人类分钟数。
正文
https://t.co/YJncE1kwkS
I think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com