论文测试 Qwen3 0.6B 至 8B 模型发现,当存储记忆过时时,模型有 92%–100% 的概率沿用旧值,且把旧笔记伪装成更新内容反而更容易骗过更大的模型。修复方式因模型规模而异:4B 和 8B 模型靠时间戳与来源元数据即可恢复大部分准确率,0.6B 和 1.7B 模型则需在输入前先解决冲突。论文建议将智能体记忆视为不可信输入。
AI agents can trust stale memory over fresh evidence, and bigger models do not reliably fix this.
Persistent memory can make an agent confidently wrong even when current evidence is available, so stale facts should be resolved before they reach the model.
The paper tests Qwen3 models from 0.6B to 8B on tasks where stored memory is outdated.
When memory is needed, they follow the stale value 92%–100% of the time.
And making an old note look newer can fool larger models even more.
Their fix depends on model size.
For 4B and 8B models, timestamps and source metadata are enough to recover most accuracy.
The 0.6B and 1.7B models need the stale conflict resolved before they see it.
So the paper suggests: treat agent memory as untrusted input.
来源:@rohanpaul_ai · x.com