跳到正文
@kimmonismus· @kimmonismus · X·· 2026-08-26AI 评分45
AI 导读

Qwen3.8-Flash-Next 今日正式发布预览版,作为即将到来的 Qwen4 架构的前瞻。被迅速删除的 ModelScope 描述显示:主模型 125B 参数、额外 51B N-gram 嵌入向量、每 token 仅激活 6B 参数,采用多模态 MoE 架构。若最终发布保持这些数字,总存储容量约 176B 参数,仅 6B 激活。

正文

Qwen may be about to drop another super interesting open-weight models of the year.

A ModelScope description that was quickly taken down reportedly revealed:

- 125B main-model parameters
- 51B additional N-gram embeddings
- Only 6B parameters active per token
- Multimodal MoE architecture

Qwen3.8-Flash-Next is officially scheduled for release today as a preview of the upcoming Qwen4 architecture.

If those numbers survive the final release, that is roughly 176B parameters of stored capacity with just 6B activated per token. Very curious to see the benchmarks.

来源:@kimmonismus · x.com