arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RP - OPSD:用于多模态大语言模型的分辨率特权在线自蒸馏

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

Qihui Zhu, Yuchen Wang, Zijian Wen, Tao Zhang, Mengjie Zhang, Yang Liu, Shuangwu Chen, Siying Wu, Jian Yang, Xiaofeng Jiang

arXiv 2607.24447首次发表:更新:

AI 中文总结

研究针对多模态大语言模型在线自蒸馏问题,利用同一图像高低分辨率视图信息差提出RP - OPSD,学生策略用低分辨率图像生成轨迹,教师策略用原始分辨率图像监督,实验表明该方法能提升性能并加快训练速度,为在线自蒸馏提供有效途径。

AI 中文摘要

在线自蒸馏(OPSD)利用教师独有的特权信息对学生生成的轨迹进行密集的令牌级监督。然而,现有方法常依赖验证过的解决方案轨迹、外部模型生成的解释或手动定位的视觉证据,限制了其在多模态大语言模型中的扩展应用。为解决此问题,我们利用同一图像高低分辨率视图间的信息差距,提出RP - OPSD。训练时,学生策略从四分之一原始分辨率的图像生成在线轨迹,教师策略用原始分辨率图像提供监督。通过最小化沿学生轨迹的输出分布差异,学生学习教师在高分辨率输入下的预测行为,增强其低分辨率能力并将学习到的改进转移到原始分辨率推理中。RP - OPSD无需额外人工标注或外部模型生成解决方案轨迹,仅需图像 - 问题对。在Qwen3.5 - 9B上的实验表明,RP - OPSD在原始分辨率下平均性能相对提高5.45%,训练速度比OPSD快1.78倍。这些结果表明分辨率差异可作为简单且可扩展的特权信息源,为多模态大语言模型的在线自蒸馏提供了有效方法。

英文摘要

On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solution traces, explanations generated by external models, or manually localized visual evidence, which limits their scalable application to multimodal large language models. To address this issue, we exploit the information gap between high- and low-resolution views of the same image and propose RP-OPSD (Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models). During training, the student policy generates on-policy trajectories from images at one-quarter of the original resolution, while the teacher policy provides supervision using the original-resolution images. By minimizing the divergence between their output distributions along the student trajectories, the student learns the predictive behavior of the teacher under high-resolution inputs, thereby strengthening its low-resolution capability and transferring the learned improvement to original-resolution inference. RP-OPSD requires neither additional human annotations nor external models to generate solution traces but only image--question pairs. Experiments on Qwen3.5-9B show that RP-OPSD achieves a 5.45\% relative improvement in average performance at the original resolution and a $1.78\times$ training speedup over OPSD. These results demonstrate that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑