跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 24 天前AI 评分49
AI 导读

微软新论文提出,长时运行智能体应在写入持久记忆前验证经验是否正确且可复用。其方案是任务结束后由独立记忆智能体以只读权限访问环境,核验后再写入长期记忆。在 CLBench 上,该设置将通过率从 39% 提升至 73%,每任务查询数从 8.8 降至 4.7,任务智能体成本从 $3.38 降至 $1.68。

正文

New Microsoft paper recommends for long-running agents, check that a lesson is correct and reusable before putting it into persistent memory.

checking agent memories against the environment before saving them made later tasks more accurate and cheaper, so verification should happen at memory-write time.

A finished agent run is not ground truth. It may contain a wrong assumption, an incomplete procedure, or a fact that becomes stale later.

Their fix is: after each task, a separate memory agent gets read-only access to the environment and checks what is worth keeping before it writes anything into long-term memory.

On CLBench, this setup raised pass rate from 39% to 73%, cut queries from 8.8 to 4.7 per task, and reduced task-agent cost from $3.38 to $1.68.

来源:@rohanpaul_ai · x.com