RSM-full 通过将相关记忆分组并整体召回,在 AMA-Bench 上约 4k prompt token 预算下保留了全量历史 83% 的质量,token 成本仅为 32%。该方法针对长时运行智能体在 prompt 预算紧张时"塞旧交互"或"检索孤立片段"两种做法的局限,主张先组织记忆再优化检索。
For memory-limited agents, organize the past before you optimize search: grouping related memories and keeping them together beat fancier retrieval under tight token budgets.
Most long-running agents either keep stuffing old interactions into the prompt or retrieve isolated chunks.
Both approaches get awkward once the prompt budget gets tight.
RSM-full takes a simpler route.
It groups related memories as they arrive, then brings those groups back together when the agent needs them.
On AMA-Bench at roughly 4k prompt tokens, it kept 83% of the quality of feeding the model its full history while using 32% of the token cost.
来源:@rohanpaul_ai · x.com