一篇关于智能体强化学习信用分配的论文提出 ProVer 方法:由 LLM 裁判对比成功与失败轨迹并定位关键片段,再采样该片段前后的续接,以成功率变化作为该步优势。相比 GRPO,ProVer 在 ALFWorld、WebShop 和 SearchQA 上对 Qwen3.5-2B 相对提升 9.91%,对 Qwen3.5-4B 提升 7.12%,且裁判模型更小时依然有效。
Good paper on credit assignment for agent RL.
The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets.
GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest.
ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage.
Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model.
Paper: https://arxiv.org/abs/2609.36178
Chat with Paper: https://academy.dair.ai/papers/targeting-pivotal-decisions-for-credit-assignment-in-agentic-reinforcement-learn-2609.36178
来源:elvis · x.com