跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-03AI 评分43
AI 导读

ContextPilot 让智能体自主决定上下文的保存、压缩与删除,并通过强化学习训练哪些上下文决策真正有效。基于 Qwen3-8B 的 ContextPilot-8B-RL 在 4 个长上下文基准上平均得分 69.40,而同一模型使用 128K 窗口且无上下文工具时仅为 45.93。

正文

Long-running agents do not just need a bigger context window. They need to learn what deserves to stay in context at all.

Long-horizon agents work better when context management is a learned policy

Long-running agents keep adding searches, tool outputs, and reasoning to the prompt. Eventually the model spends more tokens trying to separate useful facts from old clutter.

ContextPilot gives the agent control over that problem. It can plan, save important information to long-term memory, and summarize, compress, or remove old context.

Its RL training then teaches which of those context decisions actually help.

With Qwen3-8B, ContextPilot-8B-RL averaged 69.40 across 4 long-context benchmarks, versus 45.93 for the same model using a 128K window without context tools.

On BrowseComp, its context stayed around 8K–10K tokens per turn while WebExplorer-8B grew toward 30K.

Paper's overall recommendation: stop treating the entire conversation history as memory. Give agents a smaller working context they can actively manage, and train them to keep what matters.

– arxiv. org/abs/2608.28476

Title: "ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL"

来源:@rohanpaul_ai · x.com