跳到正文
@omarsar0· @omarsar0 · X·· 27 天前AI 评分46
AI 导读

KVMem 将智能体工作区溢出的上下文保留为分页 KV 状态,分布在 GPU 内存、主机内存和 NVMe 上,并用模型原生的轻量注意力空间索引选取相关历史块。在 Qwen3.8-27B 的 DeepSWE 长上下文测试中,任务成功率从 compaction 的 43.8% 提升到 48.4%。

正文

Nice paper to improve inference efficiency.

It's been a while we haven't seen good work on efficiency.

Here is why it matters:

A long-running agent's workspace outgrows its context window long before the task finishes.

The first approach commonly used, compaction, loses the fine-grained execution evidence. And text retrieval re-prefills content the model already processed.

KVMem keeps the overflow as paged KV state instead, spread across GPU memory, host memory and NVMe.

Lightweight attention-space indexes, native to the model, pick the relevant historical blocks and materialize a query-dependent view that fits inside the native context window.

On the DeepSWE long-context test with Qwen3.8-27B, task success goes from 43.8% under compaction to 48.4%.

The local deployment result stands out. It runs Qwen3.6/3.8-27B NVFP4 with MTP on a laptop with a 24GB RTX 5090, virtualizing an agent workspace up to 1M tokens, four times the model's native 256K window, at around 50 tokens per second.

Paper: https://t.co/rKIAqnzxII

来源:@omarsar0 · x.com