DeepSeek-V4.1-Flash 发布 552B 多模态 MoE 模型,KV 缓存降至每 token 890 字节
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
阅读原文
本站未展示全文,请前往来源网站阅读。
AI 导读
DeepSeek-AI 提出 DeepSeek-V4.1-Flash,一个 552B 参数的 MoE 多模态模型,支持最多一百万 token 上下文。
推荐理由
论文给出 CED 架构与 FP4 KV 缓存等设计细节,读者可据此理解长上下文部署成本的具体压缩路径。
来源:Hugging Face Daily Papers · arxiv.org