跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-27AI 评分47
AI 导读

Recuris 将智能体记忆拆分为工作记忆与经验记忆,技能选择基于当前任务状态而非完整历史,避免长程运行中的失效。在四个长程 benchmark、十个模型上,37 组模型-benchmark 组合中有 35 组任务成功率提升;tau-bench 上 GPT-5.6 Sol 提升 17.8 分、Claude Opus 5 提升 15.6 分至 87.9%。

正文

If you maintain a skill library for long-horizon agents, this one is worth your time.

(bookmark it)

It discusses one of most common topics I get asked about these days.

It shares some good ideas on how to effectively leverage memory to improve the effectiveness of long-horizon agents.

Recuris splits agent memory in two. A Working Memory tracks task progress, and an Experiential Memory holds skills.

Skill selection is grounded in the current task state instead of the full growing history, which is where long runs usually fall apart.

Because skill use is anchored to an explicit state, a failed run points at a specific memory component. A fixed Meta-Agent turns the evidence into validation-gated updates to Skill Memory, which reshape execution and produce new evidence.

Across four long-horizon benchmarks and ten models, it improves task success in 35 of 37 completed model-benchmark pairs. On tau-bench it adds 17.8 points to GPT-5.6 Sol and 15.6 points to Claude Opus 5, taking Opus 5 to 87.9 percent.

The advantage widens as the horizon grows, reaching 32.2 points on the longest tasks. Common long-horizon failures drop by up to 80 percent.

Paper: https://t.co/yx68cqRHk0

Chat with Paper: https://t.co/ivjv1GTGqX

来源:@omarsar0 · x.com