跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-26AI 评分57
AI 导读

Z.ai 正式公布此前的 Ox Alpha 实为 GLM-5.3-Flash,该模型总参数 320B,推理时仅激活 18B。相比 GLM-5.2,DeepSWE 从 46.2 升至 63.4,AutomationBench 从 26.2 升至 48.8,价格降到十分之一。架构上改用线性注意力做状态跟踪,再以稀疏注意力加轻量 indexer 检索远端上下文,注意力计算减少 3 倍、每层 KV cache 减少 4.4 倍,并引入把四个 indexer key 向量压缩为一个的 IndexPool。

正文

And its announced now officially

https://t.co/AhjbkvDczE

引用@rohanpaul_ai@rohanpaul_ai
Finally, Z .ai revealed that Ox Alpha was actually GLM-5.3-Flash. So that means over the last few days all those 100 tn tokens/day of stealth traffic capacity was running on Chinese AI chips, with tens of thousands of domestic accelerators behind the service. not an NVIDIA GPU cluster. 5.3-Flash beats GLM-5.2 at one-tenth the price with only 18B active parameters. GLM-5.3-Flash. has 320B params in total, but only 18B are active during inference. It also uses 45 layers instead of GLM-4.5's 92, cutting the amount of work required for each token. The benchmark jumps are large too: against GLM-5.2, DeepSWE rises from 46.2 to 63.4 and AutomationBench from 26.2 to 48.8. There is an architectural change as well, that cuts attention compute 3x and per-layer KV cache 4.4x versus GLM-5.3. The novelty is mainly in the combination: GLM-5.3-Flash uses linear attention for cheap state tracking, then sparse attention with a lightweight indexer to retrieve only the distant context worth revisiting, instead of repeatedly attending across the full 1M-token window. They also introduces IndexPool, which compresses four indexer key vectors into one, and says the combined design
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com