arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEGO-OPD:用于多模态在线策略蒸馏的因子化教师组合

LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

Jaeyun Shin, Hangeol Chang, Jong Chul Ye

arXiv 2610.00333首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院(KAIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LEGO-OPD,通过因子化组合语言与定位专家教师,在贝叶斯框架下独立控制视觉监督强度,实现多模态蒸馏中视觉定位与语言推理的平衡,实验验证优于现有基线。

AI 中文摘要

多模态在线策略蒸馏(OPD)旨在提升视觉定位能力,同时保持语言模型强大的推理能力。近期多教师方法结合LLM和VLM教师以提供互补的监督信号。然而,直接使用VLM的完整预测分布会将其视觉定位信号与自身的语言先验纠缠在一起,阻碍定位信息被独立传递。相反,增强视觉监督的强度可以改善感知,但可能过度强调视觉证据并削弱语言推理。为解决这一权衡,我们引入LEGO-OPD,它选择性地将语言专家和定位专家的因子组合成一个教师分布,用于多模态OPD。在广义贝叶斯框架下,语言专家为候选token提供先验,而定位专家贡献视觉似然来更新该先验,而非传递其完整预测分布。这种因子化组合使得语言推理和视觉定位可以被独立控制。我们进一步引入自适应校准,以确定在每个解码前缀处视觉似然应更新语言先验的强度。具体而言,LEGO-OPD使用定位专家的图像诱导预测偏移作为前缀相关的参考,防止视觉监督不足和过度。使用Qwen3模型的实验表明,LEGO-OPD在多模态和纯文本推理任务上均持续优于所评估的单教师和多教师OPD基线。此外,它在保持纯文本推理的同时改善了初始学生的视觉感知。

英文摘要

Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert's image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student's visual perception while preserving text-only reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑