NVIDIA 发布论文《Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills》,提出 ACES 评估框架,对同一任务分别加载与不加载目标技能各跑一次,把配对奖励差值作为 Skill Lift。
New Nvidia paper.
A high-scoring skill document is not evidence that the skill helps the agent at runtime.
Most of what an agent skill adds is discovery and workflow order, not a better final answer. So grade the trajectory, since final-answer scoring hides whether the agent ever found or followed the skill.
So run the same task twice, once with the skill loaded and once without, then compare.
NVIDIA's ACES holds task, model, harness, sandbox, grader, and supporting skills fixed, varies only whether the target skill is present, and reports the paired reward difference as Skill Lift.
On production skills with both a document score and a live run, structural and judge scores correlate with measured lift at −0.0181 and −0.0266, indistinguishable from zero.
– arxiv. org/abs/2608.20614
Title: "Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills"
来源:@rohanpaul_ai · x.com