论文发现混合Transformer模型在少量全注意力层前会产生异常巨大的内部激活值,移动这些层时巨型值随之移动,增加层数则极端值在更大网络范围内持续偏高。
Some newer efficient LLMs (using "hybrid Transformer") are cutting back on expensive attention layer of standard Transformer, and this paper shows what happens internally around the attention layers that remain.
A "hybrid Transformer" simply means a model that replaces many standard attention layers with cheaper recurrent-style layers but keeps a few full-attention layers.
That shortcut saves compute.
But this paper finds that the occasional expensive layers may have a much bigger effect on the model than their small number suggests.
Across Qwen3.5, Kimi Linear, Nemotron-H, and Zamba2, the model repeatedly produced unusually huge internal numbers immediately before those expensive look-back layers.
Move 1 of those layers, and the huge values move with it.
Use more of them, and those extreme values start staying high across larger parts of the network.
– arxiv. org/abs/2608.12149
Title: "Massive Activations in Hybrid Linear Attention LLMs: Pre-Attention Spikes and Inter-Spike Plateaus"
来源:@rohanpaul_ai · x.com