跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-31AI 评分43
AI 导读

Amazon 论文提出,KV-cache 策略不只是推理优化,还会改变模型所需的训练方式:若推理时会丢弃部分上下文,微调阶段就应采用相同的稀疏注意力策略。在 128k token 测试中,全注意力训练的模型在稀疏推理下常输出冗长无意义的回答,而策略匹配微调的模型能正常作答并停止。该方法支持任意缓存策略,可在 40 GB A100 上为 4B 模型计算梯度。

正文

New Amazon paper shows KV-cache policy is not just an inference optimization; but it can change the training regime the model needs.

If your LLM will forget parts of its context at inference, train it to forget that way too: this paper shows matching fine-tuning to the KV-cache policy can prevent long-context failures.

Sparse attention lets long-context inference use a fixed-size KV cache by keeping only part of the model’s past context. But models are often fine-tuned with full attention, then asked to work with missing memory at inference.

That mismatch can break behavior. In 128k-token tests, models trained with full attention often produced long, nonsensical answers under sparse inference, while models fine-tuned with the same cache policy learned to answer and stop normally.

The method makes this policy-matched training practical for arbitrary cache policies, including computing gradients for a 4B model on a 40 GB A100.

– arxiv. org/abs/2608.19920

Title: "Learning how to Forget: Fine-tuning for Long-Context Sparse Attention"

来源:@rohanpaul_ai · x.com