跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 21 天前AI 评分59
AI 导读

GLM-5.3 被用于优化运行自身的推理基础设施,包括生产环境内核与并发修复,使 GLM-5.3-Flash 在 10 万+ 加速器上的吞吐提升到三倍。该系统从首次运行到生产就绪用了不到两周,团队以本地正确性测试、执行轨迹、微基准和端到端测量构成密集反馈,工程师设定目标与边界,智能体负责分析、提出假设、改代码和做实验。

正文

A brilliant post from the GLM-5.3 team on RSI (recursive self-improvement).

GLM-5.3 was already used to optimize the infrastructure that runs GLM itself, including production kernel and concurrency fixes. It helped triple GLM-5.3-Flash throughput on 100,000+ accelerators.

Their early self-improvement loop: the model improves its serving system, that system runs the model, and the engineering knowledge accumulates for the next optimization cycle.

Engineers still set objectives and boundaries, while the agent handled analysis, hypotheses, code changes, and experiments.

In one test, Prefill plus KV Transfer lagged Prefill alone by over 20%, and the agent traced the slowdown to the Python GIL, and releasing that lock cut the gap below 1%

Another kernel change reached a 1.71x speedup over the prior version by eliminating repeated FP32 normalization and gating work.

引用@Zai_org@Zai_org
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://t.co/yUf6OpJD7c
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com