微软与加州大学圣地亚哥分校的论文发现,模型经 SFT 后看似更好,却可能因稀有正确行为被抹除而成为更差的 RL 起点。TailSFT 通过过滤损失已大幅下降的序列,把训练转向模型仍未充分拟合的样本,在相同 RL 设置下最终 pass@1 最高提升 3.93 个百分点。
New Microsoft plus Univ of San Diego paper shows a model can look better after SFT but actually be a worse starting point for RL, because SFT may wipe out rare correct behaviors that RL needs to discover.
TailSFT preserves more of those behaviors, and with the same RL setup it produced up to 3.93 percentage points higher final pass@1.
standard SFT can make RL harder by overtraining already-fit examples; TailSFT filters them and gives the later RL stage a better starting point.
TailSFT changes SFT by filtering sequences whose loss has already dropped most relative to the base model.
That shifts training toward examples the model still underfits, with one goal: keep correct responses reachable under repeated sampling so RL has useful behavior to reinforce.
来源:@rohanpaul_ai · x.com