AI 导读
DeepSeek 发布 DeepSeek-V4.1-Flash,552B 主干采用 causal encoder-decoder,prefill 阶段仅激活 8B 参数,decode 阶段激活 16B。该模型通过稀疏查找访问 196B Engram 记忆,并借助 bounded replay 使持久 KV 相比 V4-Flash 减少约 8 倍。
推荐理由
给出了 V4.1-Flash 的稀疏激活与记忆查找设计,读者可据此对比 V4-Flash 的持久 KV 占用变化。
正文
Congrats to @deepseek_ai on releasing DeepSeek-V4.1-Flash!
> 552B backbone, with a causal encoder-decoder activating just 8B params at prefill, 16B at decode
> 196B Engram memory accessed through sparse lookups
> ~8x less persistent KV than V4-Flash through bounded replay https://t.co/dbUG1Rutco
来源:@SemiAnalysis_ · x.com