跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-07AI 评分44
AI 导读

微软与康奈尔大学论文提出 Free Pause Tokens,让模型预测下一个 token 时获得额外计算,且不增加 token、不扩大 KV cache、不增加解码步骤。

正文

New Microsoft + Cornell Univ paper gives a method that matched a standard model trained on 50% more tokens, while adding only 14% more training time and ~1% inference latency.

Normally, the same hidden state has to do 2 jobs: keep track of the context and predict what comes next.

The paper gives prediction its own extra computation, but without adding another token, growing the KV cache, or adding another decode step.

You do not need it for the whole training run.

When the free pause was switched on after 42.5% of training, it kept about 94% of the full quality gain while training took 1.33× the wall-clock time of the normal model.

At equal node-hours, the phased versions still beat standard training.

So the recommendation is: train normally for most of the run, then add the extra prediction computation near the end.

– arxiv. org/abs/2609.03807

Title: "Free Pause Tokens"

来源:@rohanpaul_ai · x.com