arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05131cs.CVcs.AI

OPD-V:结合模态平衡的视觉在线策略自蒸馏

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态大语言模型的模态不平衡问题,提出视觉在线策略自蒸馏范式OPD-V,通过正、负教师模型实现模态平衡,在多基准与骨干上提升推理性能并降低训练成本。

中文摘要 AI 辅助

在线策略自蒸馏(OPSD)已成为提升多模态大语言模型(MLLM)视觉推理能力的标准后训练方法。现有方法从不同输入源提取特权信息以指导自蒸馏,但这些设计忽略了MLLM推理中固有的模态不平衡问题:当文本信息主导生成时,模型无法充分整合多模态输入,导致精心设计的特权信息未被充分利用,限制了OPSD的有效性。为探究该局限,我们构建了基于放大图像的正教师模型和基于掩码图像的负教师模型,二者呈现不同程度的模态不平衡;其推理正确性与token logit的变化表明,模态平衡本身可作为特权信息。基于此发现,我们提出OPD-V,一种视觉OPSD范式,通过正、负教师模型实例化该信息;正模态平衡logit间隔定义了模态平衡信任域,用于选择自蒸馏所用的在线策略token。在6个基准、4个MLLM骨干和5种后训练方法上的实验显示,OPD-V可持续提升推理性能,同时降低训练成本。

英文摘要

On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Ludwig Maximilian University of Munich(慕尼黑大学)
  • Munich Center for Machine Learning(慕尼黑机器学习中心)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑