跳到正文
@AravSrinivas· @AravSrinivas · X·· 14 天前AI 评分42
AI 导读

我们分享一项关于后训练方法的新研究,它让 Perplexity Computer 智能体通过模仿优质轨迹、并明确纠正本可避免的错误(比如错误的工具调用,即使整体轨迹是成功的)来从真实用户会话中学习。该方法结合了拒绝采样微调(RFT)与提示引导的自蒸馏,在线上 A/B 测试中将工具调用失败率降低了约 21%。

正文

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests

引用@perplexity_ai@perplexity_ai
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. https://t.co/3MFrp1yxDt
在 X 查看被引用的帖子

来源:@AravSrinivas · x.com