论文发现技能检索评估存在严重缺陷:智能体在检索到技能的任务上表现可能更差,但整体分数仍因任务难度差异而显得更好。现有评估只比较“检索 vs 未检索”任务,忽略了任务类型本身不同。论文提出 RAE 方法:当检索发生时,在同一任务上禁用技能访问重新运行并对比结果。
A skill-enabled agent can look better overall while performing worse on the exact tasks where it retrieved skills,
This paper finds a nasty evaluation failure: positive retrieval gains can hide negative performance on the very tasks where retrieval happened.
Most evaluations make a flawed comparison: they check whether tasks where the agent chose to retrieve a skill scored better than tasks where it chose not to.
But those may be completely different kinds of tasks. If the agent tends to retrieve skills on easier problems, retrieval will look helpful even if the skill itself did nothing, or even made the answer worse.
This paper’s fix, RAE, is simple: when retrieval happens, rerun that exact task without skill access and compare the result.
来源:@rohanpaul_ai · x.com