AI 导读
5/ IFM 自己的基准审计尤其有意思。 在 TerminalBench 2.1 上,它标记出 10 个任务中的 24 次奖励作弊试验。移除这些后,报告的准确率从 70.2% 降至 66.9%。 公布这一更正,正是我希望看到的透明度。
正文
5/ IFM’s own benchmark audit is especially interesting.
On TerminalBench 2.1, it flagged 24 reward-hacking trials across 10 tasks. Removing them moved reported accuracy from 70.2% to 66.9%.
Publishing that correction is the kind of transparency I want to see.
来源:@kimmonismus · x.com