跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-23AI 评分45
AI 导读

Google DeepMind 提出免训练方法 Recirculation,在推理时把激活值回灌模型,让前馈 Transformer 具备递归能力,无需重训即可像动态系统一样追踪信念状态,生成成本不变,串行计算全在 prefill 阶段。在 Gemma3 系列上,自适应版本在冻结原始权重、仅做轻量超参调优的情况下,困惑度降低 23%,GSM8k 准确率提升 21%。

正文

You don't often see one-word titles in AI papers.

That aside, strong recommend this paper from Google DeepMind.

I think this is an interesting training-free approach to evolve model architectures by leveraging the model itself to inform architectural modifications.

Something like this could also inspire even more robust recursive self-improvement approaches.

Approach details below:

A feedforward transformer can only update its internal state as many times as it has layers. Long generations need more updates than that, so chain-of-thought ends up doing basic state tracking in text.

Recirculation adds recurrence at inference time.

The model feeds activations back through itself during prefill, which lets it act like a dynamical system and track belief states without any retraining.

Generation cost stays flat. All the serial work happens in prefill.

On the Gemma3 family, the adaptive variant cuts perplexity 23% and lifts GSM8k accuracy 21%, with the original weights frozen and only light hyperparameter tuning.

Paper: https://t.co/H8LlBJG4pJ

Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

来源:@omarsar0 · x.com