跳到正文
@perplexity_ai· @perplexity_ai · X·· 14 天前AI 评分25
AI 导读

我们使用拒绝采样微调来模仿有用的步骤,并用提示引导的自蒸馏来纠正错误。 GLM 5.2 用相同的权重对同一段记录回合评分两次:一次带纠正提示,一次不带。训练让无提示的下一 token 预测与带提示时的预测对齐。

正文

We use rejection sampling fine-tuning to imitate useful steps and hint-guided self-distillation to correct errors.

GLM 5.2 scores the same recorded turn twice, using the same weights: once with a corrective hint and once without it. Training aligns the hint-free next-token predictions with those made with the hint.

来源:@perplexity_ai · x.com