跳到正文

#评测/基准

今日 1 条
今天10月2日周五
  1. Chubby♨️59

    作者引用 Epoch AI 的 Epoch Capabilities Index 数据:Claude Opus 5.5 以 167 分登顶,略高于 GPT-6 Astra;Claude Sonnet 5.5 大致追平 Claude Fable 5.1(165 分)。作者评论称 Anthropic 的中档模型在发布几个月内就达到旗舰水平,并表示很期待 Fable 5.5。榜单见 epoch.ai/eci。

    引用Epoch AI@EpochAIResearch

    Claude Opus 5.5 has taken the top spot on the Epoch Capabilities Index (ECI) with a score of 167, narrowly ahead of GPT-6 Astra. Claude Sonnet 5.5 has roughly matched Claude Fable 5.1 (165).

10月1日周四
9月22日周二
9月18日周五
  1. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个为 Verified、9 个为 Flawed、2 个信息不足暂无法评审。Karina Nguyen 转发并称此举能激励行业打造真正高质量的基准。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

9月8日周二
8月19日周三