AI 导读
在 batch 1 时,平均 TPOT 大致等于 step time 除以 accept length。我们两边都做了优化。 SGLang 缩短了每个 target step。DSpark 使用并行 draft backbone、用于局部依赖的轻量 Markov head,以及用于验证调度的 confidence head。
正文
At batch 1, mean TPOT is roughly step time divided by accept length. We worked on both.
SGLang shortened each target step. DSpark uses a parallel draft backbone, a lightweight Markov head for local dependency, and a confidence head for verification scheduling.
来源:@AntLingAGI · x.com