AI 导读
如今,仅凭通过/失败分数基本无法解读评测结果。 我在基准测试中看到的许多失败,都源于过于严格的隐藏测试,某些情况下模型的答案比预期的评测结果更合理。
正文
it's basically impossible to interpret evals by looking at just at the pass/fail scores these days
many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes more sense than the expected eval result
来源:@trq212 · x.com