一项针对自我演化编码智能体的新研究提出 EvoMal 攻击,共享技能库中的恶意技能不会被直接调用,而是被智能体当作编写模板复制、保存并执行,在库中自行扩散。在 153 个与工具相关的 SWE-bench Verified 任务、六个模型上,智能体自我投毒率为 20.3% 至 41.8%,被投毒的库最终持有的恶意技能是植入数量的 4.9 至 9.0 倍。
论文量化了共享技能库中恶意技能被智能体自行复制扩散的机制,并给出一种提示词缓解办法,可供搭建技能库的团队参考。
Important read if you build with agent skills.
Shared skill libraries are treated as a safe way for coding agents to reuse each other's work.
New research shows they propagate malware.
EvoMal plants a malicious skill in the library and never invokes it. The agent retrieves it as an authoring template, writes a new skill that preserves the payload, stores it, and runs it.
Each authored copy re-enters the library and gets imitated again.
Across six models on 153 tool-relevant SWE-bench Verified tasks, the agent self-poisoning rate runs 20.3% to 41.8%. Poisoned libraries end up holding 4.9 to 9.0 times as many malicious skills as were planted.
Deleting every planted skill does not clean it up. Qwen3 still shows 68% at round five because the agent-authored copies remain.
A counter-prompt that discourages banner-style copying drops it to 6.7% with no significant task-completion loss.
Paper: https://t.co/Njo9uJYXbN
Chat with Paper: https://t.co/da4LQ8dBKT
来源:@omarsar0 · x.com