跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-25AI 评分36
AI 导读

Google DeepMind 新论文提出 Recirculation:在推理时把深层激活的一小部分回注到浅层,无需改动权重即可让浅层看到深层已算出的上下文信息。在 Gemma3 上,10 个语言建模数据集有 9 个困惑度下降,12B 模型最多降 35%,但多选题提升有限;代价转移到无法并行的 prefill 阶段,且收益幅度无法跨模型家族迁移。

正文

New Google DeepMind paper shows, a transformer's shallow layers never see what its deeper layers have already worked out about the context. Leaking a little of it back down during inference recovers part of what was lost.

So if your model keeps losing state over a long input, the fix may not need retraining, though the cost moves to prefill, which can no longer run in parallel.

The paper shows this working on frozen weights, but the size of the benefit does not transfer across model families.

Recirculation mixes a small fraction of a deep layer's activations into a shallow layer at the next input step, weights untouched, and sets that against the off-the-shelf model.

On Gemma3, perplexity drops on 9 of 10 language-modeling datasets, by as much as 35% for the 12B model, though multiple-choice gains are modest.

– arxiv. org/abs/2608.17981

Title: "Recirculation"

来源:@rohanpaul_ai · x.com