跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-20AI 评分44
AI 导读

对自我改进智能体的攻击可以比攻击本身存活得更久:一次不安全的成功可以被转化为可复用的技能,并影响完全干净的未来任务。 而且由于基准测试会重置除已学技能文件之外的一切,任何后续危害都可以被明确追溯到智能体选择学习和保留的内容。 – https://t.co/TOFzp7Zoen 标题:"Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents"

正文

An attack on a self-improving agent can outlive the attack itself: one unsafe success can be converted into a reusable skill and influence completely clean future tasks.

And because the benchmark resets everything except the learned skill file, any later harm can be traced specifically to what the agent chose to learn and preserve.

– https://t.co/TOFzp7Zoen

Title: "Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents"

来源:@rohanpaul_ai · x.com