跳到正文
@trq212· @trq212 · X·· 25 天前AI 评分34
AI 导读

如今,仅凭通过/失败分数基本无法解读评测结果。 我在基准测试中看到的许多失败,都源于过于严格的隐藏测试,某些情况下模型的答案比预期的评测结果更合理。

正文

it's basically impossible to interpret evals by looking at just at the pass/fail scores these days

many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes more sense than the expected eval result

来源:@trq212 · x.com