AI 导读
DeepSeek 报告提出一种 Causal Encoder–Decoder 推理方案:causal encoder 处理完整 prompt 并构建 decoder global KV,decoder 只处理末尾 128 个 token,通过 SWA Bounded Replay 重建 local KV。
正文
假设你输入:
“下面是约 10 万个 token 的项目讨论记录。请总结尚未解决的决策,并解释最近一次修改。”
模型开始回答前,需要先处理这段历史。这个阶段称为 prefill。
在 Causal Encoder–Decoder 架构中,causal encoder 处理完整 prompt,并提供构建 decoder global KV 所需的representation。
按照 DeepSeek 报告中的推理方案,decoder 随后只处理 prompt 末尾的 128 个 token,通过 SWA Bounded Replay 重建 local KV,为开始生成做好准备。
收益主要体现在两个方面:减少 prefill 计算,以及通过跨层共享减少 global KV 的重复存储。
这也是报告的题目重点,极致的压榨 KV Cache Compression
Pushing the Limits of KV Cache Compression
来源:@dongxi_nlp · x.com