跳到正文
@natolambert· @natolambert · X·· 15 天前AI 评分20
AI 导读

强烈赞同。这种情况比我预期的要少得多——部分原因是构建高质量环境通常涉及相当昂贵的验证(测试强模型)。不过,还可以有更多。

正文

Strong agree. There’s way less of this than I would expect — partially due to the fact that making a high quality environment usually involves fairly expensive verification (testing strong models). Still, there can be much more. https://t.co/p9HqZgn5Kv

引用@Thom_Wolf@Thom_Wolf
releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of sharing high quality pretraining data but in the new RLVR paradigm https://t.co/Sg2mYfwswI
在 X 查看被引用的帖子

来源:@natolambert · x.com