AI 导读
Epoch AI 推出 Benchmark Reviews 基准审计计划,首批覆盖 15 个 AI 基准,其中 4 个通过验证、9 个被列为有缺陷、2 个信息不足。通过验证的包括 WeirdML v2、ExploitBench v0.1、PostTrainBench v1.1 和 SimpleQA Verified,被列为有缺陷的包括 Terminal-Bench 4.0.0、SWE-Bench Verified、SWE-Bench Pro、Humanity's Last Exam 等。Ethan Mollick 评论称这项工作有助于了解当前 AI 基准测试的状况,以及一些常用基准存在的问题。
推荐理由
Epoch AI 公布 15 个 AI 基准的审计结果,其中 9 个被列为有缺陷,涉及 SWE-Bench 等多个常用评测。
正文
Epoch continues to do some of the best public benchmarking work on AI. This is helpful (and tells us how bad the state of benchmarking is, and how terrible some of our favorite benchmarks are) https://t.co/So60XsnIMv
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review. https://t.co/xU43tv32gH在 X 查看被引用的帖子
来源:@emollick · x.com