AI 导读
有缺陷的基准测试存在一个或多个实质性缺陷,我们认为用户需要了解这些缺陷才能准确解读结果,最常见的情况是超过 20% 的任务存在影响准确率的错误。这种情况下,我们会发布一份关于所发现缺陷的有限说明。
正文
Flawed benchmarks have one or more substantive flaws we believe users need to be aware of to accurately interpret results, most commonly that >20% of the tasks have accuracy-impacting errors. In this case, we publish a limited writeup of the flaws we found.
来源:@EpochAIResearch · x.com