Salesforce 测试了两种基于记忆的智能体方法,发现 WebArena 默认任务顺序下 ReasoningBank 提升 1.5 分,打乱顺序后反而下降 4.5 分。原因是默认顺序偏向简单任务在前,智能体早期学到更干净的经验,但也会存入错误教训(如在无 API 环境中推荐 API)并反复调用,导致 71% 的情况下结果更不稳定。即使提供更好的任务细节和环境反馈,也只能挽回 31% 的下降。
The big problem with self-improving agents is that memory can compound mistakes just as easily as it compounds useful lessons.
Salesforce tested 2 memory-based agent methods and found a pretty uncomfortable pattern.
With WebArena’s default task order, ReasoningBank improved performance by 1.5 points.
Shuffle those tasks, and it dropped by 4.5 points instead.
Why? The default order tended to put easier tasks first, so the agent learned cleaner lessons early.
But memory works both ways.
Agents also saved bad lessons, like recommending APIs in an environment where APIs were impossible, then kept pulling those memories back into future tasks.
Results became more unstable in 71% of cases.
Even giving the memory system better task details and environment feedback recovered only 31% of the drop.
So “learning from experience” is only useful if the agent is learning the right thing.
– arxiv. org/abs/2608.18066
Title: "On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification"
来源:@rohanpaul_ai · x.com