发表机构
Nankai University; Northwestern Polytechnical University; Institute of Automation, Chinese Academy of Sciences; NKIARI(南开大学; 西北工业大学; 中国科学院自动化研究所; NKIARI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出OraRL算法,将标注作为神谕回滚解耦优势估计,实现高效可扩展的视频MLLM强化学习,在多项视频感知基准上超越现有模型,解码速度大幅提升。
AI 中文摘要
多模态大语言模型(MLLMs)已成为统一视频感知的主流范式。然而,在大型多任务数据集上进行后训练仍具挑战性,因为现有强化学习方法即使采用成本高昂的思维链(CoT)生成,也仅能采样到包含少量高质量回滚的在线策略组。本文研究视频多模态大语言模型强化学习后训练的样本效率与可扩展性,提出OraRL。我们发现标注被忽视的作用:除了对回滚进行评分外,每个标注都可作为“神谕回滚”进入其在线策略组,成为直接的正优化目标。但直接整合神谕并非易事:高奖励神谕会提升组基线,反转原本正向的策略优势,我们将此失败称为“优势反转”。OraRL的核心是解耦优势估计器:策略回滚确定无基准的基线,而神谕-策略差距同时调节方向增益与独立的神谕优势。符号平衡剪枝提升效率:仅保留神谕及每种符号下最强的回滚,OraRL仅需SFT的2.2倍步长时间,远低于带CoT的GRPO所需的4.9倍。OraRL随模型规模与数据量扩展,在0.8B到9B的主干模型上表现优于其主干,在最多10万提示词场景下优于GRPO。无需思维链时,Video-ORA-9B解码时间为130ms,而非4780ms。与各自先前最优模型相比,它将时间mIoU从62.5提升至66.0,跟踪AO从73.0提升至78.2,分割性能从64.3提升至70.4,三项基准的空间智能宏平均从51.0提升至56.1;在VSI-Bench上,它得分为73.1,而GPT-5为55.0、Gemini-3-Pro为55.1。
英文摘要
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as existing reinforcement learning methods sample on-policy groups with few high-quality rollouts even with costly chain-of-thought (CoT) generation. In this paper, we study the sample efficiency and scalability of RL post-training for video MLLMs and introduce OraRL. We identify an overlooked role for annotations: Beyond scoring rollouts, each can enter its on-policy group as an oracle rollout, a direct positive optimization target. Direct oracle integration, however, is nontrivial: a high-reward oracle raises the group baseline and inverts otherwise positive policy advantages, a failure we term advantage inversion. At the core of OraRL is a decoupled advantage estimator: policy rollouts determine an oracle-free baseline, while the oracle-policy gap modulates both a directional gain and a separate detached oracle advantage. Sign-balanced pruning improves efficiency: by retaining only the oracle and the strongest rollouts of each sign, OraRL requires just 2.2x the step time of SFT, less than half the 4.9x required by GRPO with CoT. OraRL scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts. Without chain-of-thought, Video-ORA-9B decodes in 130 ms instead of 4,780 ms. Compared with the respective prior best models, it raises temporal mIoU from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4, and the three-benchmark spatial-intelligence macro average from 51.0 to 56.1; on VSI-Bench, it scores 73.1 against 55.0 for GPT-5 and 55.1 for Gemini-3-Pro.
CommentsProject page: https://orarl.github.io/