跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-02AI 评分37
AI 导读

微软新论文发现,若目标是降低推理内存,免训练的滑动窗口注意力(SWA)优于多数需后训练的线性注意力改造方案,只需保留近期小窗口加前 4 个 sink token。

正文

So happy to see this new Microsoft paper.

If lower inference memory is the goal, this paper finds training-free Sliding Window Attention beats most retrofitted linear-attention methods, making it the simpler default to try first.

Keep only a small recent window, plus the first 4 “sink” tokens that models rely on.

With a 64-token window, this training-free setup had the best average downstream score in 9 of 11 model comparisons and recovered 99.0% of the full-attention baseline average.

Many linear-attention alternatives need additional post-training; this version of SWA needs none.

The gap grew on long-context reasoning.

At 4K context, SWA reached 17.2%–23.0% on the Needle-in-a-Haystack tasks, while LoLCATs reached at most 5.8%; on BABILong, SWA scored 15% versus 3%.

In their speed and memory test, the 64-token SWA setup was fastest and used the least memory.

Full attention still wins badly on long context, but for fixed, low memory without retraining, the paper recommends trying SWA with attention sinks first.

– arxiv. org/abs/2608.28444

Title: "Sliding-window beats linear attention"

来源:@rohanpaul_ai · x.com