AI 导读
我们将公开的 DSpark 配方适配到 Ling,使用了分布对齐的数据、架构消融实验,以及一个感知接受率的损失函数。 SplitServe Trainer 将一台 8 GPU 节点拆分用于 draft 训练和 SGLang 目标推理,使长上下文在线训练保持在本地。https://t.co/guVczj3qoO
正文
We adapted the public DSpark recipe to Ling with distribution aligned data, architecture ablations, and an acceptance aware loss.
SplitServe Trainer divided one 8 GPU node between draft training and SGLang target inference, keeping long context online training local. https://t.co/guVczj3qoO
来源:@AntLingAGI · x.com