AI 导读
小米 MiMo 团队(罗福莉)分享 MiMo-V2.5 系列 API 降价背后的推理系统重构。该系列基于 Hybrid Sliding Window Attention 架构,KVCache 存储可压缩到全注意力的约 1/7;团队重新设计 KVCache 管理、层级缓存和 prefix-cache tree,并优化调度策略与 Prefill/Decode 流水线。
正文
最近大家看到小米的MiMo 模型的降价!
我今天看了一下用了120 w 差不多花了3块多钱!
正好看到小米MiMo团队罗福莉分享的一篇技术博客。
V2.5系列刚把API价格降下来,背后其实是他们把推理系统彻底重构了一遍。
他们用的Hybrid Sliding Window Attention架构,能把KVCache存储压缩到全注意力的约1/7。
但罗福莉他们很清楚,架构优势在真实生产流量里不会自动变现。
于是团队重新设计了KVCache管理、层级缓存和prefix-cache tree,针对SWA特有的缓存难题做了专项处理,同时深度优化了调度策略和Prefill/Decode流水线。
在真实生产流量验证后,有效KVCache容量提升了接近5倍,主流框架下的服务端缓存命中率稳定在93%到95%。
再叠加MoE配置调优和多模态推理优化,才真正把长上下文推理成本打下来,支撑了这次降价。
这恰巧说明,好架构只是天花板,把它真正落地成可规模化、低成本的生产能力,才是决定模型性价比的关键。
Inference Optimizations Behind the MiMo-V2.5 Series API Price Reductions Read the full technical blog: mimo.xiaomi.com/blog/mimo-v2… The V2.5 model family, including MiMo-V2.5 and MiMo-V2.5-Pro, is built on a Hybrid Sliding Window Attention (Hybrid SWA) architecture, which compresses KVCache storage to roughly 1/7 that of Full Attention. However, architectural advantages rarely translate directly into measurable gains in production serving. To realize these gains, we redesigned KVCache management, tiered caching, and the prefix-cache tree; addressed key challenges in SWA KVCache handling; and optimized scheduling as well as the Prefill/Decode pipeline. Validated on real production traffic, these optimizations have increased effective KVCache capacity by nearly 5x, with server-side cache hit rates averaging 93%–95% across mainstream harness frameworks. Together with MoE configuration tuning and multimodal inference optimizations, they enable more efficient long-context inference and form part of what makes the recent API price cuts possible.在 X 查看被引用的帖子
来源:@berryxia · x.com