跳到正文
@emollick· @emollick · X·· 28 天前AI 评分33
AI 导读

@METR_Evals 的长时程(long horizons)指标是否实际上已经饱和了? 自 5 月以来,那张最著名的 AI 进展图就没有更新过,但其他已发表的工作表明,在 harness 中,Fable(我猜 GPT-6 也一样)实际上能完成 18+ 周的工作。https://t.co/2AjjXcQVCg

正文

Is the @METR_Evals long horizons measure effectively saturated?

No updates to the most famous graph of AI progress since May, but other published work suggests you can effectively get 18+ weeks of work out of Fable (and I assume GPT-6) in harnesses. https://t.co/2AjjXcQVCg

来源:@emollick · x.com