Zai 发布 GLM-5.3 Flash(Ox Alpha),320B MoE 每 token 仅激活 18B 参数,采用 MIT 开源许可,原生多模态并支持 1M 上下文。
原文列出六项基准数据与 MIT 开源、1M 上下文等规格,便于读者判断这一稀疏 MoE 的效率定位。
GLM-5.3 Flash ("Ox Alpha") official: Benchmarks attached. This looks exceptional for its size!
GLM-5.3-Flash might be one of the most impressive efficiency releases yet.
It is a 320B MoE with only 18B parameters active per token, yet Zai reports:
- 84.3 on Terminal-Bench 2.1, nearly matching Claude Opus 4.8 at 85.0
- 63.4 on DeepSWE, ahead of Opus 4.8 and DeepSeek V4 Vision Exp
- 48.8 on AutomationBench, ahead of Opus 4.8 and GPT-5.6 Terra
- The highest GDPval-AA v2 score in its comparisonIt also beats the much larger GLM-5.2 across all six reported benchmarks while costing one-tenth as much to serve.
Open weights, MIT licensed, natively multimodal, 1M context.
Important caveat: 18B active parameters does not make it a normal local 18B model. All 320B weights still need to be stored. But in terms of intelligence per active parameter, this looks exceptional!
The upcoming Ox Alpha is GLM-5.3 Flash (as expected): 320b total parameters, 18b active. Outperforming GLM-5.2 at 1/10th of its price and approaching Opus 4.8 on coding and agentic benchmarks. Big things incoming! https://t.co/xZZmOD1Ghu https://t.co/H2fZ9idgVa在 X 查看被引用的帖子
来源:@kimmonismus · x.com