AI 导读
@METR_Evals 的长时程(long horizons)指标是否实际上已经饱和了? 自 5 月以来,那张最著名的 AI 进展图就没有更新过,但其他已发表的工作表明,在 harness 中,Fable(我猜 GPT-6 也一样)实际上能完成 18+ 周的工作。https://t.co/2AjjXcQVCg
正文
Is the @METR_Evals long horizons measure effectively saturated?
No updates to the most famous graph of AI progress since May, but other published work suggests you can effectively get 18+ weeks of work out of Fable (and I assume GPT-6) in harnesses. https://t.co/2AjjXcQVCg
来源:@emollick · x.com