跳到正文
原文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-26精选AI 评分66
AI 导读

Z.ai 披露此前匿名上线的模型 Ox Alpha 实为 GLM-5.3-Flash,作者转述的报道称该服务由数万个国产加速器支撑,而非 NVIDIA GPU 集群。

推荐理由

原文确认 Ox Alpha 即 GLM-5.3-Flash,并给出参数规模、架构改动与基准对比,便于对照其相对 GLM-5.2 的位置。

正文

Finally, Z .ai revealed that Ox Alpha was actually GLM-5.3-Flash.
So that means over the last few days all those 100 tn tokens/day of stealth traffic capacity was running on Chinese AI chips, with tens of thousands of domestic accelerators behind the service. not an NVIDIA GPU cluster.

5.3-Flash beats GLM-5.2 at one-tenth the price with only 18B active parameters.

GLM-5.3-Flash. has 320B params in total, but only 18B are active during inference.

It also uses 45 layers instead of GLM-4.5's 92, cutting the amount of work required for each token.

The benchmark jumps are large too: against GLM-5.2, DeepSWE rises from 46.2 to 63.4 and AutomationBench from 26.2 to 48.8.

There is an architectural change as well, that cuts attention compute 3x and per-layer KV cache 4.4x versus GLM-5.3.

The novelty is mainly in the combination: GLM-5.3-Flash uses linear attention for cheap state tracking, then sparse attention with a lightweight indexer to retrieve only the distant context worth revisiting, instead of repeatedly attending across the full 1M-token window. They also introduces IndexPool, which compresses four indexer key vectors into one, and says the combined design

引用@rohanpaul_ai@rohanpaul_ai
OxAlpha is a new iteration of GLM, from China’s Z .ai And it will change how you run long running agent fast. --- bloomberg .com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek https://t.co/ZOTlnBgOsy
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com