跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-26AI 评分57
AI 导读

一篇论文在 SWE-bench Verified 上做受控实验,发现智能体基准分数中 harness 引起的方差是模型引起的 7.8 倍,9 组模型对比里有 6 组排名因所用 harness 不同而翻转。

正文

Great paper on why agent leaderboard comparisons are hard to trust.

It's on the hot topic of how much of an agent benchmark score actually belongs to the harness.

The harness is the layer between the model and the task. It builds the context the model sees, mediates tool calls, validates outputs, and decides when to retry or stop. Every score comes out of a model and a harness together, but only the model gets reported.

The authors ran a controlled grid to measure this. Three frontier models, three harness configurations, 100 tasks from SWE-bench Verified, with task order, execution environment, step budget, and evaluation script all held fixed.

Swapping the harness moved GLM-5.1 by 13.0 points. Swapping the model inside a fixed harness moved scores by 3.0, 2.5, and 5.0 points.

Harness-induced variance came out 7.8x larger than model-induced variance, and 6 of 9 model-pair comparisons flipped their ranking depending on which harness ran.

Public leaderboards show the same thing. On SWE-bench Verified Mini, HAL reports a 34 point swing for Claude Sonnet 4.5 across scaffolds and nearly 48 points for o4-mini.

They propose a Harness Card, a structured disclosure across seven layers, so you can tell whether a score gap came from the model, the harness, or the interaction.

Paper: https://t.co/pAE8edBsB9

Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

来源:@omarsar0 · x.com