跳到正文
Hugging Face Daily Papers·· 2026-08-21AI 评分43

FlashPrefill V2:面向长上下文 LLM 服务的块稀疏 Prefill 注意力

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

阅读原文

本站未展示全文,请前往来源网站阅读。

AI 导读

FlashPrefill V2 将块稀疏 Prefill 注意力从算法原型推进到可部署的长上下文 LLM 服务方案,新增均值修正项抑制近似误差,并重写稀疏注意力算子以对齐 FlashAttention-3/4、支持 FP8 推理。

来源:Hugging Face Daily Papers · arxiv.org