跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 16 天前AI 评分42
AI 导读

智能体工作是 Step 5 Preview 落后于同等 Intelligence Index 分数模型的地方。在 GDPval-AA 上它落后于 GLM-5.3(max)和 Qwen3.8 Max(1,566 Elo 对 1,646 和 1,668),在 AA-Briefcase 上(1,432 对 1,525 和 1,640),在 Terminal-Bench 4.0 上(33% 对 42% 和 39%)。在知识和推理方面则相反,它在 Humanity's Last Exam(46% 对 42% 和 43%)、CritPt(21% 对 19% 和 18%)和 AA-Omniscience Index(16 对 14 和 12)上均领先两者。

正文

Agentic work is where Step 5 Preview lags models at a similar Intelligence Index score. It sits behind GLM-5.3 (max) and Qwen3.8 Max on GDPval-AA (1,566 Elo against 1,646 and 1,668), AA-Briefcase (1,432 against 1,525 and 1,640) and Terminal-Bench 4.0 (33% against 42% and 39%). The reverse holds on knowledge and reasoning, where it is ahead of both on Humanity's Last Exam (46% vs 42% and 43%), CritPt (21% vs 19% and 18%) and the AA-Omniscience Index (16 vs 14 and 12)

来源:@ArtificialAnlys · x.com