跳到正文
原文
@omarsar0· @omarsar0 · X·· 2026-08-26精选AI 评分76
AI 导读

阿里 Qwen 团队开放 Qwen3.8-Flash 权重,该模型为多模态 MoE,总参数 125B、每 token 激活 6B,并带 51B N-gram 嵌入,官方称其为 Qwen4 架构的早期预览。生产版本将上线 QwenCloud API,输入 $0.16/1M tokens、输出 $0.47/1M tokens,官方还给出 DeepSWE 1.1 58.7、SWE-bench Pro 62.5 等成绩。作者 @omarsar0 认为随附的技术报告比发布本身更值得读。

推荐理由

发布信息列出了 Qwen3.8-Flash 的参数量、激活规模与定价,可作为判断高效多模态 MoE 路线的具体参照。

正文

What's better than an open-weight multimodal model release?

Well, the technical report. I just love how these labs like Qwen and DeepSeek continue to drop gem after gem.

Qwen3.8-Flash is the latest in efficient multimodal MoE models. Worth reading the report. https://t.co/Z73tEodbgo https://t.co/q6NzzKJTwA

引用@Alibaba_Qwen@Alibaba_Qwen
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG
在 X 查看被引用的帖子

来源:@omarsar0 · x.com