跳到正文
@MiniMax_AI· @minimax_ai · X·· 2026-05-06AI 评分53
AI 导读

MiniMax-M2.7 现已在 Artificial Analysis 的六家推理服务商上线,SambaNova 以 435 output tokens/s 领先,约为次快的 Fireworks(127 output tokens/s)的 3.4 倍,MiniMax 一方 API 为 57 output tokens/s,Novita 54、GMI 41、Together AI 29。价格上 Fireworks 约 $0.22/1M tokens 混合价与 MiniMax 一方 API 持平但快约 2.2 倍,SambaNova 价格约为其他提供商的 2-3.5 倍;Fireworks、MiniMax、Novita、Together AI 提供 80% 缓存命中折扣,GMI 与 SambaNova 无此折扣。MiniMax 转发时表示,M2.7 达到 400 TPS 后延迟几乎无法察觉。

正文

Thanks to the @SambaNovaAI team. Once M2.7 hits 400 TPS, latency becomes virtually imperceptible.

引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys
MiniMax-M2.7 is now available across six inference providers on Artificial Analysis, with significant differentiation in speed and price @SambaNovaAI leads on speed at 435 output tokens/s, >3x faster than any other provider. @FireworksAI_HQ, @novita_labs, @togethercompute, and @GMI_cloud have all matched @MiniMax_AI's first-party API pricing, while SambaNova is 2x higher. Key takeaways: ➤ Fireworks and SambaNova are on the Pareto frontier for Speed vs. Price. At 127 output tokens/s and ~$0.22 per 1M tokens blended, Fireworks is ~2.2x faster than MiniMax's first-party API at the same blended price, whereas SambaNova delivers 435 output tokens/s but at ~2-3.5x the blended price of the other providers (depending on cache usage) ➤ SambaNova is the fastest provider at 435 output tokens/s, ~3.4x the next fastest provider (Fireworks at 127 output tokens/s). The remaining providers run substantially slower: MiniMax’s first-party API at 57 output tokens/s, Novita at 54, GMI at 41, and Together AI at 29 ➤ Cache discounts vary across providers. Fireworks, MiniMax, Novita, and Together AI offer 80% cache hit discounts, while GMI and SambaNova do not offer a discount. For cache-heavy workloads, this can materially increase the relative pricing for GMI and SambaNova ➤ Optimal provider choice depends on workload. SambaNova may be more suited to latency-sensitive deployments, albeit at a higher cost, while Fireworks may be more suitable for high-volume workloads that are not as latency-sensitive
在 X 查看被引用的帖子

来源:@MiniMax_AI · x.com