arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10321cs.CL

面向视觉-语言模型适配的在线策略蒸馏:低质量多模态数据上的有效范式

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

  • The University of Hong Kong(香港大学)
  • Great Wall Motor(长城汽车)
  • Wuhan University(武汉大学)
  • China University of Petroleum (East China)(中国石油大学(华东))
  • University of Science and Technology of China(中国科学技术大学)
  • Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongyuan Zhang, Xianda Guo, Yanlun Peng, Qianlong Yang, Yubin Guo, Pinhan Fu, Mulin Chen, Xiaozhen Qiao, Ping Luo

AI总结:

提出在线策略蒸馏框架OnPoKD,将蒸馏目标构建视为动态策略决策,通过轻量控制器自适应平衡教师监督、零样本先验和硬标签,提升视觉-语言模型在低质量多模态数据上的适配与泛化性能。

AI中文摘要:

知识蒸馏提供了一条高效途径,可将任务适配的视觉-语言教师模型迁移至紧凑的学生模型。当前视觉-语言蒸馏方法中的训练目标通常由教师预测构建,并统一应用于所有训练样本,这在类别偏移和领域偏移下变得不可靠。本文认为,蒸馏目标的构建应被视为一种动态的训练决策,而非固定配方。为此,我们提出了OnPoKD,一种用于视觉-语言模型适配的在线策略蒸馏框架。据我们所知,OnPoKD是首个将在线策略蒸馏应用于视觉-语言模型适配的框架,其通过将目标构建学习为策略决策来实现。OnPoKD学习一个轻量级控制器,利用来自教师模型、学生模型和零样本先验的可靠性线索与分歧线索,构建样本自适应的目标。控制器不依赖固定的教师预测,而是通过有界策略动作动态平衡教师监督、零样本先验引导和硬标签锚定,使蒸馏目标能够适应不同的样本可靠性和训练阶段。策略控制器通过验证反馈进行更新,鼓励目标构建优化可迁移性,而不仅仅是拟合训练分布。由于控制器仅在训练期间使用,OnPoKD可以无缝集成到现有的视觉-语言蒸馏流程中,同时保持原始推理架构和测试时成本。在基础到新类泛化和跨数据集迁移基准上的大量实验表明,OnPoKD持续优于强视觉-语言蒸馏基线。

英文摘要:

Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all training samples, making it unreliable under class and domain shifts. In this paper, we argue that distillation target construction should be treated as a dynamic training decision rather than a fixed recipe. To this end, we propose OnPoKD, an on-policy distillation framework for vision-language model adaptation. To the best of our knowledge, OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision. OnPoKD learns a lightweight controller that constructs sample-wise adaptive targets using reliability and disagreement cues from the teacher model, student model, and zero-shot prior. Instead of relying on a fixed teacher prediction, the controller dynamically balances teacher supervision, zero-shot prior guidance, and hard-label anchoring through bounded policy actions, allowing the distillation target to adapt to varying sample reliability and training stages. The policy controller is updated with validation feedback, encouraging target construction to optimize transferability rather than merely fitting the training distribution. Since the controller is only used during training, OnPoKD can be seamlessly integrated into existing vision-language distillation pipelines while preserving the original inference architecture and test-time cost. Extensive experiments on Base-to-novel generalization and Cross-dataset transfer benchmarks show that OnPoKD consistently improves over strong vision-language distillation baselines.

↑