NVIDIA 新论文指出,企业技能库常用的结构扫描评分与 LLM 评审质量的相关性仅为 Spearman rho 0.14(基于 145 个真实技能)。论文提出 ACES 的 Skill Lift 方法:同一模型、沙箱、工作区和评分器下,同一任务分别加载与不加载技能各跑一次,比较智能体完成差异。
Very interesting new paper from NVIDIA.
(bookmark it)
It takes a closer look at evaluating agent skills.
Enterprise teams are starting to leverage shared skill libraries, and the review gate is typically a scanner that checks structure, style, and security.
NVIDIA measured whether that gate predicts anything.
Across 145 real skills from internal and public catalogs, structural scan scores correlate with LLM-judge quality at a Spearman rho of 0.14.
ACES proposes Skill Lift instead.
In other words, run the same task twice under the same model, sandbox, workspace, and scorer, once with the skill loaded and once without. Then you measure the difference in what the agent completed.
They scored 947 paired cases from 58 production skills across four harnesses, normalizing trajectories into a shared Agent Trajectory Interchange Format, so results compare across harnesses.
They fins that the largest process-metric gains appear in skill execution, behavior check, and skill efficiency.
Paper: https://t.co/NDIO8yF3nr
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com