跳到正文
原文
@kimmonismus· @kimmonismus · X·· 2026-08-26精选AI 评分72
AI 导读

Qwen3.8-Flash-Next 与 GLM-5.3-Flash 两款开放权重模型发布,均为 MoE 架构,每 token 分别激活 6B 和 18B 参数。

推荐理由

推文列出两款开放权重模型的参数规模与多项基准结果,可供读者对比其与前沿闭源模型的差距。

正文

Today is a huge day for open-weight and local AI.

Qwen3.8-Flash-Next is a 125B MoE with another 51B n-gram embedding parameters, yet only 6B are active per token.

It beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks, including SWE-bench Pro, CoWorkBench, GPQA Diamond and LiveCodeBench.

GLM-5.3-Flash is even larger: 320B total, 18B active. It scores 84.3 on Terminal-Bench 2.1 versus 85.0 for Opus 4.8, while beating Opus on DeepSWE, AutomationBench and GDPval-AA v2.

MIT licensed, natively multimodal, 1M context.

One important distinction: 6B or 18B active parameters does not mean 6B or 18B storage. The complete weights still need to be stored.

So “local” here means a serious workstation or local server, not an ordinary laptop.

The fact that we now have models operating at the level of Opus 4.6 - or even Opus 4.8 - a level of performance that was state of the art only a few months ago and can now, at least theoretically, be run locally, should serve as a wake-up call, especially for U.S. frontier labs.

Not only because the gap continues to narrow, but also because there appears to be enormous demand for open AI. In that sense, what has been released here is genuinely welcome.

So yeah, great day for everyone!

来源:@kimmonismus · x.com