AI 导读
论文《Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents》指出,长周期智能体安全不能简化为反复运行安全的短周期轨迹。攻击者可将恶意证据拆分到多个单独看似无害的步骤中,仅监控单条轨迹的监测器无法获得足够上下文来区分攻击与正常工作。
正文
If an autonomous agent remembers across iterations, its safety system cannot afford to forget between them..
long-horizon agent safety cannot be reduced to repeatedly running a safe short-horizon trajectory.
An attacker can split malicious evidence across several individually benign-looking steps, so a trajectory-only monitor never sees enough context to distinguish the attack from normal work.
– arxiv. org/abs/2608.27141
Title: "Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents"
来源:@rohanpaul_ai · x.com