斯坦福一篇论文提出 Prefix Sliding,在推理生成过程中丢弃推理链中间 token,只保留含指令与工具的 prefix 和最近数千 token 的窗口,无需任何训练即可让现有模型运行快 3 倍,性能与 full attention 持平。
Banger paper from Stanford on efficient test-time scaling.
If you run agents that think for a long time, this one is worth your time.
(bookmark it)
Long reasoning keeps the entire trace in memory through full attention.
This means that the hardest problems, the ones that need the most thinking, are also the ones that cost the most to run.
The authors measured what the middle of a reasoning trace is actually worth.
Intermediate tokens steadily lose importance as the model keeps going.
Their new approach, Prefix Sliding, drops those tokens. It keeps the prefix, which holds the instructions and the available tools, plus a window of the last few thousand tokens. Everything in between gets discarded during generation.
Total memory stays capped no matter how long the model reasons.
Without any training, this runs existing models 3x faster while matching full-attention performance, and it enables RL rollouts past 100,000 tokens.
Paper: https://t.co/HzwSZ7fCdh
Chat with Paper: https://t.co/OfjtVjamIC
来源:@omarsar0 · x.com