一篇论文发现上下文压缩会在少数压缩事件中损害长程智能体的任务完成与可靠性,典型情况是丢掉未解决的任务条件(如 Venvo 任务中被删的 only from coworkers 过滤条件)或把已读 API 说明压成模糊文字,导致智能体重读文档、多花约五步。
Context compression is a huge bottleneck for long-running agents.
This work finds that context compression hurts long-horizon agents at a few specific points.
They propose PAIR, which replays the agent from the same state with and without a given compression, instead of comparing whole runs that differ in many random ways.
A typical compression adds a few extra steps. The large drops in success come from a small number of compression events.
The harmful compressions drop task conditions the agent hasn't resolved yet. In one Venmo task, the summary dropped the "only from coworkers" filter and reported the total of all 36 payments as the answer.
Other compressions reduce API specs the agent already read to vague prose, so the agent reopens the docs and logs in again, which adds about five steps.
PAIR then diagnoses what information those compressions dropped and rewrites the matching sections of the compression prompt. The agent, compressor model, and tools stay fixed.
On AppWorld, OfficeBench, and tau-Bench Retail, it gives the most consistent task completion of any compressed method and comes close to running with no compression at all.
Compression also lowers run-to-run reliability before it makes tasks unsolvable, so check consistency across repeated runs in your own evals.
Paper: https://arxiv.org/abs/2609.36526
Chat with Paper: https://academy.dair.ai/papers/adapting-context-compression-for-long-horizon-agents-with-counterfactual-continu-2609.36526
来源:DAIR.AI · x.com