斯坦福新论文发现,在 3 个企业级 Agent benchmark 上,不到 3% 的分数差异来自 Agent 本身,7–23% 来自 Agent 与具体任务的交互,排行榜排名高未必适合你的工作流。
New Stanford paper shows a leaderboard can tell you which agent scored highest on that benchmark, but it may not tell you which agent is actually better for your work.
Across 3 enterprise agent benchmarks, this paper finds that less than 3% of score variation comes from the agent itself.
A much larger 7–23% comes from the interaction between the agent and the specific task.
In plain English: an agent that looks better overall may simply fit the benchmark's task mix better.
The problem gets worse on hard tasks. On τ2-bench action checks, reliability falls from 0.752 overall to 0.000 on the hardest task quartile.
And across 50 train/test splits, projected reliability correlated with held-out reliability at r = -0.90.
So a leaderboard can look precise while being a weak guide for your actual workflow.
So for choosing agents, the better question is not "Which agent ranks 1st?" but "Which agent works reliably on my task types, difficulty level, and cost constraints?"
That also supports routing different task classes to different agents instead of forcing 1 model to win everything.
– arxiv. org/abs/2608.11323
Title: "Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations"
来源:@rohanpaul_ai · x.com