AI 导读
拒绝采样微调只模仿成功的会话。这会强化模型从中恢复的错误,并丢弃失败会话中的证据。 我们模仿成功会话中的有用步骤,并用两种会话类型中经过验证的提示来纠正错误。https://t.co/KXS1jgnzyB
正文
Rejection sampling fine-tuning imitates only successful sessions. This can reinforce errors the model recovered from and discard evidence in failed sessions.
We imitate useful steps from successful sessions and correct errors with validated hints in either session type. https://t.co/KXS1jgnzyB
来源:@perplexity_ai · x.com