AI 导读
第一阶段支持 TP 和 PP,单节点与多节点部署,非量化权重,以及 block-wise FP8。 下一步:DP 和 EP、更多格式、KV cache 恢复、跨 GPU 共享、CUDA graph 序列化,以及持久化 kernel 预热。 冷恢复目标为 10 秒以内。🚀
正文
Phase 1 supports TP and PP, single node and multi node deployment, unquantized weights, and block-wise FP8.
Next: DP and EP, more formats, KV cache restore, cross GPU sharing, CUDA graph serialization, and persistent kernel warmup.
The cold recovery target is under 10 seconds. 🚀
来源:@AntLingAGI · x.com