AI 导读
Hugging Face 把 OpenAI 的数学题转化为开源 RL 环境,发布在 Hugging Face 数据集 FineEnvs/openai-math 上。该版本尚处早期粗糙阶段,验证器只接受完全一致的原始形式化,等价证明仍可能得 0 分。Clément Delangue 认为,把开放研究变成人人可构建的开源可执行环境,是训练更好开源模型的重要方向。
正文
We turned @OpenAI's math problems into open-source RL environments on @huggingface!
Early and rough (the verifier only accepts the exact original formalization, so an equivalent proof can still score 0), but this is what it looks like when a research release becomes executable training infrastructure for everyone.
This is an exciting direction imo: turning open research into open executable environments that anyone can build on to train better open models!
https://huggingface.co/datasets/FineEnvs/openai-math
来源:clem 🤗 · x.com