跳到正文
原文
Epoch AI· @EpochAIResearch · X·· 15 天前AI 评分50
AI 导读

Epoch AI 推出 Benchmark Reviews,用于审计 AI 基准,首批发布 15 个基准的评审结果。其中 4 个 Verified,包括 WeirdML v2、ExploitBench v0.1、PostTrainBench v1.1 和 SimpleQA Verified;9 个 Flawed,包括 Terminal-Bench 4.0.0、SWE-Bench Verified、SWE-Bench Pro、Humanity's Last Exam、DeepSWE v1.1、TextQuests、Lech Mazur Writing、BFCL v4 和 HealthBench Professional;另有 CritPt 和 FrontierCode 因信息不足暂无法评审。

正文

Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

来源:Epoch AI · x.com