DeepSeek 发布 V4.1-Flash,距离 7 月的 V4-Flash 更新约六周,改用 Causal Encoder–Decoder 新架构并具备原生视觉理解能力。
距上代仅六周,DeepSeek 在新架构下同时给出参数量、KV-cache 占用与 API 价格的变化,可与前代直接对照。
DeepSeek just released V4.1-Flash with a new architecture, six weeks after its July V4-Flash update.
July’s release improved post-training while keeping the architecture unchanged. (Same with GLM-5.3/Flash) V4.1 introduces a Causal Encoder–Decoder architecture with native visual understanding.
DeepSeek reports:
- 552B MoE parameters, with 8B active during input processing and 16B during output generation.
- KV-cache requirements cut to ¼ of the HBM and ⅛ of the SSD storage versus the previous generation.
- Lower API prices.
These are *significant* jumps in just a few weeks with post training.
This is the new reality we have to adapt to: weekly releases with significant improvements.
The company says Flash now beats V4-Pro on capability, cost and speed. Starting September 14, V4-Pro API requests will temporarily route to V4.1-Flash until V4.1-Pro arrives.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6在 X 查看被引用的帖子
来源:@kimmonismus · x.com