跳到正文
@natolambert· @natolambert · X·· 23 天前AI 评分31
AI 导读

强化学习扩展的一个基本思路:我们能否把更多算力分配给更难的问题? 我们这样做了:如果你的 GRPO 组里所有回答都是错的,就以概率 P(约 0.9)多采样一些——以寻找更多梯度非零的 GRPO 批次。它有效! 叫做 "Never Give Up" https://t.co/bNPn6EQsiE https://t.co/VcE8I1kKtE

正文

An basic idea in scaling RL: Can we allocate more compute to the harder problems?

We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works!

Called "Never Give Up" https://t.co/bNPn6EQsiE https://t.co/VcE8I1kKtE

来源:@natolambert · x.com