跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分38
AI 导读

弗吉尼亚理工新论文提出 Hybrid Latent Attention,将较旧的 token 缓存为紧凑向量供注意力直接读取,使循环 LLM 在每块 GPU 上可容纳的请求量提升 4.0 至 8.8 倍,且精度损失很小。

正文

New Virginia Tech paper shows how cache older tokens as compact vectors that attention reads directly, and a looped LLM fits 4.0 to 8.8 times as many requests per GPU at little accuracy cost.

Looped models run each token through the same layers several times, and every pass adds to the KV cache, so fewer requests fit on a GPU.

Hybrid Latent Attention keeps the last 128 tokens exact and squeezes older ones into small vectors that every loop reads without rebuilding anything.

On Ouro models, throughput rose up to 7.4 times at 16K tokens, and accuracy stayed above 97% of the original on math, knowledge and reasoning.

Gains are largest on long contexts. You can retrofit an existing looped checkpoint by training only the new parts.

– arxiv. org/abs/2610.07940

Title: "Hybrid Latent Attention for Looped Language Models"

来源:Rohan Paul · x.com