IBM 论文提出 BenchDrift,一个对现有基准做审计的框架,通过在不改变答案的前提下系统改写同一问题,测量模型分数因措辞产生的漂移。在 8 个模型和 3 个基准上,最好与最差准确率的差距平均达 74.7 个百分点,且在基线准确率高于 60% 的模型-基准组合中,改写打破的正确答案多于挽回的;即使最高置信度分组,18.5% 的正确答案在保持语义的改写后丢失。
A model can know the answer and still fail because you asked the same question differently.
New IBM paper introduces BenchDrift, an auditing framework for existing benchmarks, that systematically rephrases the same questions without changing their answers, then measures how much a model's score moves purely because of wording.
Across 8 models and 3 benchmarks, the gap between best-case and worst-case accuracy averaged 74.7 percentage points.
Stronger models were actually more exposed.
For every model-benchmark pair above 60% baseline accuracy, rephrasing broke more correct answers than it recovered.
And confidence did not solve this.
Even in the highest-confidence bucket, 18.5% of correct answers were lost after meaning-preserving rewording.
So if your model selection depends on small score differences, test several equivalent phrasings or report the range.
– arxiv. org/abs/2608.11694
Title: "The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance"
来源:@rohanpaul_ai · x.com