跳到正文
@TencentHunyuan· @tencenthunyuan · X·· 2026-06-09AI 评分45
AI 导读

🚀推出 UniRL,一个面向统一多模态模型的 RL 基础设施。同时带来两个新的 RL 算法:DRPO 和 Flow-DPPO。 一个 RL 循环覆盖扩散/流匹配模型、LLM/VLM 以及统一多模态模型👇 代码:github.com/Tencent-Hunyuan/U… (没错——U(你)-ni-(需要)RL 😉)

正文

🚀Introducing UniRL, an RL infra for unified multimodal models. Together with two new RL algorithms: DRPO and Flow-DPPO.

One RL loop across diffusion/flow matching models, LLMs/VLMs, and unified multimodal models👇

Code: github.com/Tencent-Hunyuan/U…

(yes — U(you)-ni-(need) RL 😉)

来源:@TencentHunyuan · x.com