一篇论文在固定模型和 3,679 道题的前提下,只改动提示词格式、选项顺序和打分方式等常规评测选择,发现结果变化很大。gemma4-31b 的得分在 31% 到 89% 之间波动,12 个模型中有 4 个在至少一种有效配置下排到第一。相邻模型平均差距的 95.7% 来自答案会随评测配置改变的题目,最大的不稳定来源是打分方式,即生成答案还是选择最高似然选项。
LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the leaderboard.
The paper keeps the models and 3,679 questions fixed, then changes only ordinary evaluation choices such as prompt format, option order, and scoring method.
Those choices move the results a lot.
gemma4-31b scores anywhere from 31% to 89%, and 4 of the 12 models reach rank 1 under at least 1 valid setup.
Even more telling, 95.7% of the average gap between neighboring models comes from questions whose answers change when the evaluation setup changes.
The biggest source of instability is how answers are scored: generating an answer versus choosing the highest-likelihood option.
来源:@rohanpaul_ai · x.com