Claude Opus 5.5 has taken the top spot on the Epoch Capabilities Index (ECI) with a score of 167, narrowly ahead of GPT-6 Astra. Claude Sonnet 5.5 has roughly matched Claude Fable 5.1 (165).
#评测/基准
#评测/基准
今日 1 条
Chubby♨️@kimmonismusAI 评分5959
引用Epoch AI@EpochAIResearch
gabriel@gabriel1AI 评分2020
SemiAnalysis@SemiAnalysis_AI 评分4848
Yuchen Jin@Yuchenj_UWAI 评分3030Google 回来了??? 全面优于 Astra 和 Opus 5.5。 如果这不只是刷榜,我很想看到他们重新加入竞赛。

Logan Kilpatrick@OfficialLoganKAI 评分3434如果你在用 AI 构建产品,你应该花超过 25% 的时间做基准测试,并努力让模型实验室关注这些基准测试 这是加速公司进展的最简单路径
Karina@karinanguyenAI 评分5656引用Epoch AI@EpochAIResearchIntroducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
jietang@jietangAI 评分1616更新了 AA index……现在 Fable 5.1 和 GPT-astra 的……一样了……

Karina@karinanguyenAI 评分3737Grok 4.6 在 DiligenceBench 金融测试中以约 52–53% 位列第二,与 Claude Opus 5 基本持平,Sonnet 5 以 46.2% 落后。
