研究对比同一批历史轨迹的两种形式发现,蒸馏成 SKILL.md 的技能版本比保留更多执行细节的 Workflow Memory 表现高 6.06 个百分点,说明智能体并非获得更多经验,而是同样的经验被更好地打包。轨迹分析显示 65.7% 的技能生效案例靠流程锚定,仅 4.5% 靠补充缺失知识,因此技能主要作用于执行层面;但用错场景或过于死板执行时,技能反而会拖累表现。
Agent skills work for a very specific reason: they turn messy past experience into a clean procedure the agent can follow.
The researchers gave agents the same past trajectories in 2 forms: Workflow Memory, which keeps more execution detail, and a distilled SKILL.md.
The skill version performed 6.06 percentage points better than Workflow Memory.
Because the agent was not getting more experience. It was getting the same experience packaged better.
Their trajectory analysis makes the mechanism clearer: 65.7% of skill cases worked through procedural anchoring, while only 4.5% worked by supplying missing knowledge.
So skills mainly help with execution: what to do first, which tools to use, what to verify, and which mistakes to avoid.
This also explains the failure mode. A skill can still hurt when it is used in the wrong situation or followed too rigidly.
Overall takeaway, self-improving agents need better distillation and application of experience, not just bigger memory libraries.
– arxiv. org/abs/2608.14036
Title: "Demystifying Agent Skills: Why They Work-Until They Don't"
来源:@rohanpaul_ai · x.com