arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03141cs.CV

视频多模态大语言模型后训练的行为包优化

Behavior Pack Optimization for Video MLLM Post-Training

  • Peking University(北京大学)
  • The University of Melbourne(墨尔本大学)
  • Stanford University(斯坦福大学)
  • Chengdu Minto Tech(成都明图科技)

机构由 AI 辅助整理,请以论文原文为准。

Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou

AI总结:

针对视频MLLM后训练中依赖外观和语言先验而忽视时间证据的问题,提出行为包优化(BPO),通过跨反事实视角联合评分并采用锚定相对优势,在多个基准上显著提升准确率与弃权能力。

AI中文摘要:

视频多模态大语言模型(MLLMs)在视频问答基准上持续攀升,然而,打乱帧顺序、掩盖支持答案的片段或遮挡目标对象几乎不会改变它们的预测。其准确性依赖于外观和语言先验,而非问题所要求的时间证据。我们将此归因于后训练的单位:奖励是在对原始片段的单一响应上计算的,因此模型从未被要求在不同视角下保持一致的行为。我们提出行为包优化(BPO),该方法用跨反事实视角(由问题类型选择)的输出行为包替代单一响应,并联合评分。行为包奖励要求:当干预无关时保持稳定性,当关键证据被移除时保持敏感性,当没有证据剩余时弃权(不执行)。为了在小行为包大小下保持该目标的稳定性,BPO使用锚定相对优势:原始视角上的响应作为每个提示的参考,而不是混合视角上的组均值。在TempCompass、MVBench和NExT-QA上,BPO将Qwen2.5-VL-7B-Instruct的宏平均准确率提高了4.7个百分点,时间困难子集提高了7.8个百分点,弃权F1分数提高了20.0个百分点,相较于来自相同SFT检查点的预算匹配的vanilla GRPO基线。这些提升迁移到Video-MME、LongVideoBench以及LLaVA-Video-7B;消融实验证实这些提升源于视角集,而非滚动计数。我们希望这种行为包级别的视角能为视频MLLM和多模态后训练社区提供一个有用的起点,因为该领域正朝着基于证据的视频推理方向发展。

英文摘要:

Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence the question asks for. We trace this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. We propose Behavior Pack Optimization (BPO), which replaces the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly. The pack reward asks for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. To keep this objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. On TempCompass, MVBench, and NExT-QA, BPO improves the macro accuracy of Qwen2.5-VL-7B-Instruct by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint. The gains transfer to Video-MME, LongVideoBench, and to LLaVA-Video-7B; ablations confirm they follow the view sets, not the rollout count. We hope this pack-level perspective offers a useful starting point for the video MLLM and multimodal post-training community as the field moves toward evidence-grounded video reasoning.

补充信息

↑