DeepSeek 发布 V4.1-Flash,作者引用 DeepSeek 的测试称其在多项编码与智能体基准上追平或超过 GPT-5.6 Sol,而输出 token 价格低 94%。
文中给出 DeepSeek V4.1-Flash 的定价与 DeepSWE 基准对比,读者可据此判断价格战对智能体调用成本的影响。
Whats really exciting about DeepSeek 4.1 Flash its insane efficiency and pricing.
It matches or beats GPT-5.6 Sol on several coding and agent benchmarks, according to DeepSeek’s tests, at a fraction of the token price (-94%!)
For example: on DeepSWE, a software engineering benchmark:
- Previous V4-Flash: 54.4%
- New V4.1-Flash: 74.2%
- GPT-5.6 Sol: 73.0%
DeepSeek also cut prices versus V4-Flash: 32% for fresh input, 57% for cached input and 9% for output.
Even at peak rates, Flash charges $1.20 per million output tokens versus Sol’s $20. That is 94% less. Off-peak, it drops to $0.60.
A new architecture makes reading large inputs cheaper, while sharing and compressing stored context reduces memory requirements. Particularly useful for agents repeatedly reading code, documents and tool results.
DeepSeek is truly back in the price war. This is the real moat.
DeepSeek just released V4.1-Flash with a new architecture, six weeks after its July V4-Flash update. July’s release improved post-training while keeping the architecture unchanged. (Same with GLM-5.3/Flash) V4.1 introduces a Causal Encoder–Decoder architecture with native visual understanding. DeepSeek reports: - 552B MoE parameters, with 8B active during input processing and 16B during output generation. - KV-cache requirements cut to ¼ of the HBM and ⅛ of the SSD storage versus the previous generation. - Lower API prices. These are *significant* jumps in just a few weeks with post training. This is the new reality we have to adapt to: weekly releases with significant improvements. The company says Flash now beats V4-Pro on capability, cost and speed. Starting September 14, V4-Pro API requests will temporarily route to V4.1-Flash until V4.1-Pro arrives.在 X 查看被引用的帖子
来源:@kimmonismus · x.com