跳到正文
@berryxia· @berryxia · X·· 2026-05-30AI 评分56
AI 导读

小米 MiMo 团队(罗福莉)分享 MiMo-V2.5 系列 API 降价背后的推理系统重构。该系列基于 Hybrid Sliding Window Attention 架构,KVCache 存储可压缩到全注意力的约 1/7;团队重新设计 KVCache 管理、层级缓存和 prefix-cache tree,并优化调度策略与 Prefill/Decode 流水线。

正文

最近大家看到小米的MiMo 模型的降价!

我今天看了一下用了120 w 差不多花了3块多钱!

正好看到小米MiMo团队罗福莉分享的一篇技术博客。

V2.5系列刚把API价格降下来,背后其实是他们把推理系统彻底重构了一遍。

他们用的Hybrid Sliding Window Attention架构,能把KVCache存储压缩到全注意力的约1/7。

但罗福莉他们很清楚,架构优势在真实生产流量里不会自动变现。

于是团队重新设计了KVCache管理、层级缓存和prefix-cache tree,针对SWA特有的缓存难题做了专项处理,同时深度优化了调度策略和Prefill/Decode流水线。

在真实生产流量验证后,有效KVCache容量提升了接近5倍,主流框架下的服务端缓存命中率稳定在93%到95%。

再叠加MoE配置调优和多模态推理优化,才真正把长上下文推理成本打下来,支撑了这次降价。

这恰巧说明,好架构只是天花板,把它真正落地成可规模化、低成本的生产能力,才是决定模型性价比的关键。

引用Fuli Luo (@_LuoFuli)@_LuoFuli
Inference Optimizations Behind the MiMo-V2.5 Series API Price Reductions Read the full technical blog: mimo.xiaomi.com/blog/mimo-v2… The V2.5 model family, including MiMo-V2.5 and MiMo-V2.5-Pro, is built on a Hybrid Sliding Window Attention (Hybrid SWA) architecture, which compresses KVCache storage to roughly 1/7 that of Full Attention. However, architectural advantages rarely translate directly into measurable gains in production serving. To realize these gains, we redesigned KVCache management, tiered caching, and the prefix-cache tree; addressed key challenges in SWA KVCache handling; and optimized scheduling as well as the Prefill/Decode pipeline. Validated on real production traffic, these optimizations have increased effective KVCache capacity by nearly 5x, with server-side cache hit rates averaging 93%–95% across mainstream harness frameworks. Together with MoE configuration tuning and multimodal inference optimizations, they enable more efficient long-context inference and form part of what makes the recent API price cuts possible.
在 X 查看被引用的帖子

来源:@berryxia · x.com