Rohan Paul 认为「时间跨度」正成为衡量智能体能力更实用的指标,并引用 OpenAI 数据指出任务越长、无需人工干预的成功率越低,从 15 分钟以内的 86% 降至 64-128 小时区间的约 16%。他指出模型可以在局部表现很强,但在长执行轨迹上仍可能不可靠。他由此提出,或许可以转向类似「每完成一个任务消耗多少人工分钟」这样的基准。
I think "time horizon" is becoming one of the more useful ways to talk about agent capability.
At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common.
A model can be extremely capable locally and still be unreliable over a long execution trajectory.
Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene?
maybe we are moving towards a benchmark, something like
human minutes consumed per completed task.
OpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com