跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-18AI 评分40
AI 导读

论文发现混合Transformer模型在少量全注意力层前会产生异常巨大的内部激活值,移动这些层时巨型值随之移动,增加层数则极端值在更大网络范围内持续偏高。

正文

Some newer efficient LLMs (using "hybrid Transformer") are cutting back on expensive attention layer of standard Transformer, and this paper shows what happens internally around the attention layers that remain.

A "hybrid Transformer" simply means a model that replaces many standard attention layers with cheaper recurrent-style layers but keeps a few full-attention layers.

That shortcut saves compute.

But this paper finds that the occasional expensive layers may have a much bigger effect on the model than their small number suggests.

Across Qwen3.5, Kimi Linear, Nemotron-H, and Zamba2, the model repeatedly produced unusually huge internal numbers immediately before those expensive look-back layers.

Move 1 of those layers, and the huge values move with it.

Use more of them, and those extreme values start staying high across larger parts of the network.

– arxiv. org/abs/2608.12149

Title: "Massive Activations in Hybrid Linear Attention LLMs: Pre-Attention Spikes and Inter-Spike Plateaus"

来源:@rohanpaul_ai · x.com