AI 导读
Duke 大学 PhD Fred Peng 团队提出 REPR-ALIGN,通过将扩散语言模型(DLM)的 hidden states 逐层用余弦相似度对齐到冻结的同架构自回归 teacher 模型,实现最高 4 倍训练加速,低数据场景下效果尤为明显。
正文
兄弟们,训练Diffusion LLM原来可以这么省?
大家都知道扩散语言模型(DLM)很香:支持双向生成、非顺序解码、灵活编辑。
但从零训一个,成本高得离谱。
Duke大学PhD Fred Peng(@pengzhangzhi1)和团队直接给出了一个反直觉的答案:
别重训了,直接对齐就行。
论文标题叫《Don’t Retrain, Align》。
核心思路很简单:
我们已经有强大的预训练Autoregressive LM(AR LM),里面已经学好了绝大部分语言表示。
DLM真正需要改的只是生成顺序和去噪行为。
所以他们提出了REPR-ALIGN:在做masked diffusion训练的同时,逐层用余弦相似度,把DLM的hidden states对齐到冻结的AR teacher模型上。
不需要加adapter,不需要改架构,只改attention mask。
结果:在他们的实验设置里,训练速度最高提升4倍,低数据场景下效果尤其明显。
一句话总结:
不要把表示空间从头重训一遍,对齐它,让模型只去重新学习解码路径就够了。
Paper:arxiv.org/abs/2605.06885
Code:github.com/pengzhangzhi/Open…
如果你在搞扩散模型、生成式AI或者长上下文生成,这篇值得立刻读。
How to Train Diffusion LLM more efficiently? Our paper has an answer for you: Don’t Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment Diffusion language models are becoming increasingly attractive: they support bidirectional generation, non-sequential decoding, and flexible editing. But training them from scratch is expensive. So a natural question is: If we already have strong pretrained autoregressive LMs, do we really need to relearn all language representations for diffusion LMs? We argue: probably not. Our view is that AR→DLM conversion should not be treated as learning language from scratch again. Much of the semantic structure is already inside the AR model. What changes is the generation order and denoising behavior. So instead of only continuing denoising training, we explicitly preserve the representation geometry of the AR model. We introduce REPR-ALIGN: during masked diffusion training, we align the hidden states of the DLM to a frozen AR teacher of the same architecture, layer by layer, using cosine similarity. No adapters. No architectural changes beyond the attention mask. Just representation alignment + masked denoising. The result: up to 4× training acceleration in our setting, with especially strong gains in low-data regimes. The main takeaway is simple: Don’t retrain the representation space from scratch. Align it, and let the model relearn the decoding path. Paper: arxiv.org/abs/2605.06885 Code: github.com/pengzhangzhi/Open… Work done with an amazing undergrad @alexisfox and advisors @Anru_Zhang @AlexanderTong7在 X 查看被引用的帖子
来源:@berryxia · x.com