跳到正文
@EpochAIResearch· @EpochAIResearch · X·· 21 天前AI 评分14
AI 导读

有缺陷的基准测试存在一个或多个实质性缺陷,我们认为用户需要了解这些缺陷才能准确解读结果,最常见的情况是超过 20% 的任务存在影响准确率的错误。这种情况下,我们会发布一份关于所发现缺陷的有限说明。

正文

Flawed benchmarks have one or more substantive flaws we believe users need to be aware of to accurately interpret results, most commonly that >20% of the tasks have accuracy-impacting errors. In this case, we publish a limited writeup of the flaws we found.

来源:@EpochAIResearch · x.com