跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-29AI 评分41
AI 导读

这篇论文是对长周期AI的一次残酷现实检验。给一个智能体一年的相互关联决策、延迟反馈以及自身过往行为的后果,它的表现相对人类就会崩溃。 研究人员测试了八款领先模型,包括GPT-5.6 Sol和Claude Opus 4.8。然而表现最好的配置——Qwen3.7-Max搭配Hermes——最终赚到的钱仅为人类参与者平均水平的27.3%。 一个完成一年期任务时仅有人类四分之一表现的系统,距离可靠的长周期执行还差得远。 https://t.co/ki6xniZg9c

正文

https://t.co/ki6xniZg9c

引用@rohanpaul_ai@rohanpaul_ai
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans. The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant. A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com