跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-09-05AI 评分39
AI 导读

Artificial Analysis 的 AA-Briefcase 评测显示,Anthropic 的 Claude Fable 5.1 和 Opus 5 位居榜首,GPT-6 Astra 与 Muse Spark 1.3 紧随其后;GPT-6 Astra 较 GPT-5.6 Sol 提升约 85 Elo 分。

正文

Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points.

AA-Briefcase is our frontier in-house evaluation with a private held-out test set. The evaluation tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

来源:@ArtificialAnlys · x.com