跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-01AI 评分44
AI 导读

一种名为Memory-Augmented Compression的免训练方法,将已解示例蒸馏为可复用推理记忆,按查询检索并注入提示词,让模型以更短推理作答,把部分计算从自回归解码转移到并行的prefill阶段。在Qwen2.5-7B上,为Chain-of-Draft加入记忆后,GSM8K准确率回升21.4分、MATH回升28.0分,延迟仍比标准CoT快1.49倍和1.14倍。

正文

This paper shows a different way to make chain-of-thought cheaper: move reusable reasoning from generation into the prompt.

Instead of asking an LLM to regenerate every reasoning step, give it the relevant reasoning pattern upfront and let it think shorter.

Memory-Augmented Compression is training-free: it distills solved examples into reusable reasoning memories, retrieves relevant ones for each query, and injects them before compressed reasoning.

That shifts some work from slow autoregressive decoding to the more parallel prefill stage.

With Qwen2.5-7B, adding memory to Chain-of-Draft recovered 21.4 accuracy points on GSM8K and 28.0 on MATH, while model latency stayed 1.49× and 1.14× faster than standard CoT.

来源:@rohanpaul_ai · x.com