跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-07AI 评分56
AI 导读

Rohan Paul 认为「时间跨度」正成为衡量智能体能力更实用的指标,并引用 OpenAI 数据指出任务越长、无需人工干预的成功率越低,从 15 分钟以内的 86% 降至 64-128 小时区间的约 16%。他指出模型可以在局部表现很强,但在长执行轨迹上仍可能不可靠。他由此提出,或许可以转向类似「每完成一个任务消耗多少人工分钟」这样的基准。

正文

I think "time horizon" is becoming one of the more useful ways to talk about agent capability.

At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common.

A model can be extremely capable locally and still be unreliable over a long execution trajectory.

Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene?

maybe we are moving towards a benchmark, something like

human minutes consumed per completed task.

引用@rohanpaul_ai@rohanpaul_ai
OpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com