跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分63
AI 导读

一项关于个人 Agent 自我改进极限的研究发现,记忆笔记只在一定范围内有帮助,之后额外内容会让 Agent 更容易违反规则。用 Claude Haiku 4.5 的实验中,违规率从无记忆时的 77% 降到 10 行记忆时的 20%,更长记忆后回升到约 25%;对持续统计的支出总额,书面规则失败率 44%,而用代码追踪失败率 0%。

正文

More memory does not keep helping personal agents, because relevant notes help up to a point and then extra lines start making the agent miss rules.

A personal agent can't learn every user preference through memory notes, and more notes eventually make it worse, so use code for anything it must count or track and keep memory short.

Written rules work for style, like signing texts with the user's first name. For a running spending total, a stated rule still failed 44% of the time, while code that kept the total failed 0%.

Memory size has a sweet spot. With Claude Haiku 4.5, violations dropped from 77% with no memory to 20% at 10 lines, then rose to about 25% with longer memories.

Agents that rewrite their own memory from user complaints improve early, then stall. The best methods ended near 48% violations, against 7.1% when the agent was simply told every preference.

来源:Rohan Paul · x.com