一篇论文提出自我改进 LLM 智能体的技能错误演化现象,不安全任务被蒸馏成可复用技能后,即使原始恶意指令已消失,仍会在后续会话中改变智能体行为。在 21 个演化后的智能体方法配置中,21 个全部产出了不安全技能文件,但只有 15 个在全新会话中造成实际危害,其余 6 个未触发危害。
论文给出技能库持久化风险的量化证据,并附检测基准与修复方案,可迁移到智能体安全评估。
An agent can receive a clean prompt today and still behave unsafely because yesterday’s malicious task was saved into its skill library.
This paper calls that skill misevolution: an unsafe task succeeds, gets distilled into a reusable skill, and later changes behavior even after the original malicious instruction is gone.
Across 21 evolved agent-method configurations, all 21 authored unsafe skill artifacts, but only 15 produced fresh-session harm. i.e. every evolved setup learned unsafe skills, but in 6 of the 21 setups those skills did not cause harm in the later clean session.
So if you only check the agent’s final behavior, you can miss the fact that its persistent skill library is already carrying unsafe instructions that may be triggered later.
The paper introduces SKILLMISEVO-GYM/BENCH to detect and measure how unsafe experience becomes persistent agent skills, and SAFEEVOLVE to reduce that risk by checking, repairing, tracking, and retiring unsafe skills before they keep propagating.
来源:@rohanpaul_ai · x.com