AI 导读
核心问题是信用分配。 一次长时程 rollout 可能耗时数十小时,而单一的终端奖励必须归因到数千次交互上。 这正是 GRPO 这类无价值(value-free)RL 方法开始吃力的地方。
正文
The core problem is credit assignment.
A long-horizon rollout may take tens of hours, while a single terminal reward must be attributed across thousands of interactions.
This is where value-free RL methods such as GRPO begin to struggle.
来源:@omarsar0 · x.com