AI 导读
🚀推出 UniRL,一个面向统一多模态模型的 RL 基础设施。同时带来两个新的 RL 算法:DRPO 和 Flow-DPPO。 一个 RL 循环覆盖扩散/流匹配模型、LLM/VLM 以及统一多模态模型👇 代码:github.com/Tencent-Hunyuan/U… (没错——U(你)-ni-(需要)RL 😉)
正文
🚀Introducing UniRL, an RL infra for unified multimodal models. Together with two new RL algorithms: DRPO and Flow-DPPO.
One RL loop across diffusion/flow matching models, LLMs/VLMs, and unified multimodal models👇
Code: github.com/Tencent-Hunyuan/U…
(yes — U(you)-ni-(need) RL 😉)
来源:@TencentHunyuan · x.com