卡内基梅隆大学与牛津大学论文提出"looped flows",让模型通过反复改进隐状态而非生成更长思维链来"思考更久",额外迭代步数可显著提升准确率。该方法针对循环模型长循环训练不稳定的问题,将每次更新训练为小型去噪任务,同时保证隐状态对下一次更新仍有用,使模型在推理时持续改进同一内部表示,意味着测试时计算不必等同于生成更多 token。
New Carnegie Mellon + Oxford Univ paper shows a model can "think longer" by repeatedly improving hidden state rather than writing a longer chain-of-thought.
and looped flows show those extra steps can materially raise accuracy.
that small models can reason better by refining hidden state for more steps, so test-time compute does not have to mean generating more tokens.
The problem is that recurrent models are hard to train over long loops: the model may keep updating its hidden state, but those updates can become unstable or stop helping.
Looped flows fix this by training each update on a small denoising task while making sure the hidden state remains useful for the next update.
This lets the model keep improving the same internal representation at inference time.
来源:@rohanpaul_ai · x.com