跳到正文
Hugging Face Daily Papers·· 2026-08-26AI 评分42

OraRL:面向视频 MLLM 的高效可扩展强化学习后训练方法

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

阅读原文

本站未展示全文,请前往来源网站阅读。

AI 导读

研究者提出 OraRL,将标注作为 oracle rollout 直接纳入 on-policy 组,并用解耦优势估计器解决由此产生的 advantage inversion 问题。

来源:Hugging Face Daily Papers · arxiv.org