Qwen3.8-Flash-Next 发布,采用 125B MoE 参数加 51B N-gram embeddings,每 token 仅激活 6B 参数。
原文给出四项架构改动与 1/9 训练成本的对比,读者可以了解高稀疏 MoE 如何压低单 token 计算量。
A bit more tl;dr about the model: Qwen3.8-Flash-Next combines 125B MoE parameters with 51B N-gram embeddings, while activating only 6B parameters per token.
The model introduces four major changes:
– Gated DeltaNet + Qwen Sparse Attention
– Gated Residual connections
– N-gram embeddings
– A Muon-based training recipe
Qwen says it required just one-ninth (1/9!!) of Qwen3.7-Plus’s training cost, while scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro and 73.9 on CoWorkBench.
Its new architecture combines Gated DeltaNet with sparse attention that selects small context blocks instead of attending across the entire sequence. Qwen reports up to 7.6× faster prefill and 4.9× faster decoding at a one-million-token context length.
so in sum: Qwen’s early preview of Qwen4 shows where model scaling is heading: more parameters, far less compute per token.
At this point, I’m just as excited about Chinese open-source releases as I am about new OpenAI models.
Qwen 3.8 Flash-Next official released: A 6B-active open model just beat Claude Opus 4.6 Max across 8 of 9 comparable benchmarks! Qwen3.8-Flash-Next is a highly sparse MoE: • 125B model parameters • 51B additional n-gram embeddings • Only 6B parameters active per token It scores: • 62.5 SWE-bench Pro • 81.0 SWE-bench Multilingual • 73.9 CoworkBench • 55.7 JobBench • 73.5 Toolathlon • 81.3 IFBench • 91.7 GPQA Diamond • 91.9 LiveCodeBench It also outperforms Qwen3.8-27B and DeepSeek-V4-Flash across most of the table. Super cool release!!在 X 查看被引用的帖子
来源:@kimmonismus · x.com