跳到正文
@TencentHunyuan· @tencenthunyuan · X·· 2026-06-09AI 评分54
AI 导读

腾讯混元发布 UniRL,一套面向统一多模态模型的 RL 基础设施,并开源两个新算法 DRPO 和 Flow-DPPO,代码已在 GitHub 放出。

正文

🚀Introducing UniRL, an RL infra for unified multimodal models. Together with two new RL algorithms: DRPO and Flow-DPPO.

One RL loop across diffusion/flow matching models, LLMs/VLMs, and unified multimodal models👇

Code: github.com/Tencent-Hunyuan/U…

(yes — U(you)-ni-(need) RL 😉)

1、Most RL stacks are built for one modality. UniRL applies a single post-training loop — generate → score → advantage → update → sync — across model families. Model and algorithm are two independent axes, so your coverage is the model × algorithm product, not a fixed recipe menu.

2、One loop, every modality: text→image, text/image→video, vision-language, text-only LLM and VLM, the LLM→diffusion prompt-enhancer, and unified autoregressive+diffusion generation (Hunyuan-Image 3 and Bagel) — a model class no single-purpose RL repo can even express.

3、Built to scale: pluggable rollout engines (train-side / SGLang / vLLM-Omni) behind one typed contract, FSDP2 sharding, and three deployment modes from a single config knob.

4、Two team-original algorithms headline the release:

FlowDPPO: Policy optimization for flow/diffusion models with trust-region masks based on exact divergence (See our paper: Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models github.com/Tencent-Hunyuan/U…)

DRPO: LLM RL with a smooth, advantage-weighted quadratic regularizer

(See our paper: Rethinking the Divergence Regularization in LLM RL [arxiv.org/abs/2606.09821])

来源:@TencentHunyuan · x.com