@Alibaba_Qwen· @Alibaba_Qwen · X·· 2026-08-26精选AI 评分75
AI 导读
通义千问发布多模态 MoE 模型 Qwen3.8-Flash 并开放权重,QwenCloud 上的 API 同步上线。模型为 125B 参数加 51B N-gram embeddings,每 token 仅激活 6B,采用 GDN + QSA 混合注意力、Gated Residual、N-gram Embedding 与 Muon 优化器,官方称其为 Qwen4 架构的早期预览,训练成本仅为 Qwen3.7-Plus 的 1/9。原生上下文 262K 可经 YaRN 扩展到 1M,评测得分 DeepSWE 1.1 58.7、SWE-bench Pro 62.5、CoWorkBench 73.9、AndroidWorld 84.5、MathVision 95.7,官方同时开放了 Qwen3.8-Flash-Next 的权重。
推荐理由
训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。
正文
API is live on QwenCloud: https://t.co/nz3wibv0Lv
🙌Let's build something! https://t.co/uQQ5m5lOXx
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG在 X 查看被引用的帖子
来源:@Alibaba_Qwen · x.com