现代AI智能体将经验积累在harness(提示词、记忆、工具、技能、路由规则)中而非模型权重里,更新任一harness组件都可能导致原有可靠行为失效,论文将此定义为harness级遗忘并给出度量方法。
Banger paper on harness continual learning.
(bookmark it)
If you already are allowing your agents to rewrite their own prompts, skills, or memory files, this one is worth your time.
(bookmark it)
Continual learning has always tracked what changes in the weights. Modern agents accumulate experience in the harness instead, across prompts, memories, tools, skills, and routing rules.
What this means is that if you update any harness component, previously reliable behavior can break with the model completely untouched. The paper names that harness-level forgetting and provides a way to measure it.
Guarded harness evolution separates proposing an update from committing it. A Continual Optimizer drafts a candidate harness from post-execution feedback, and a Continual Evaluator commits only after checking current improvement, historical retention, and validity.
Relative gains exceed 10% across textual reasoning, multimodal perception, and open-world interaction.
Paper: https://t.co/58HvPcpANi
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com