跳到正文
@TencentHunyuan· @tencenthunyuan · X·· 2026-06-09AI 评分63
AI 导读

腾讯混元发布 UniRL,一个用同一套后训练循环(generate→score→advantage→update→sync)覆盖多种模态的强化学习框架,涵盖 text→image、text/image→video、视觉语言、纯文本 LLM 与 VLM,以及 Hunyuan-Image 3 和 Bagel 这类自回归加扩散统一生成。

正文

1、Most RL stacks are built for one modality. UniRL applies a single post-training loop — generate → score → advantage → update → sync — across model families. Model and algorithm are two independent axes, so your coverage is the model × algorithm product, not a fixed recipe menu.

2、One loop, every modality: text→image, text/image→video, vision-language, text-only LLM and VLM, the LLM→diffusion prompt-enhancer, and unified autoregressive+diffusion generation (Hunyuan-Image 3 and Bagel) — a model class no single-purpose RL repo can even express.

3、Built to scale: pluggable rollout engines (train-side / SGLang / vLLM-Omni) behind one typed contract, FSDP2 sharding, and three deployment modes from a single config knob.

4、Two team-original algorithms headline the release:

FlowDPPO: Policy optimization for flow/diffusion models with trust-region masks based on exact divergence (See our paper: Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models github.com/Tencent-Hunyuan/U…)

DRPO: LLM RL with a smooth, advantage-weighted quadratic regularizer

(See our paper: Rethinking the Divergence Regularization in LLM RL [arxiv.org/abs/2606.09821])

来源:@TencentHunyuan · x.com