发表机构
The Hong Kong University of Science and Technology (Guangzhou); ByteDance Inc.; Zhejiang University of Technology; Hong Kong Baptist University(香港科技大学(广州); 字节跳动公司; 浙江工业大学; 香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉语言模型无外部监督难以改进的问题,提出NOPD自蒸馏方法,利用干净与损坏输入预测差异产生自监督信号,在多视觉推理任务中有效,能提升多个模型在多个基准测试中的表现。
AI 中文摘要
训练后可使视觉语言模型(VLM)理解人类指令并执行各种下游任务。当前训练后方法通常依赖人工标注数据、外部模型蒸馏、有人类反馈的强化学习或可验证答案,限制了其在无外部监督下的改进能力。为此提出NOPD(有噪学生策略性自蒸馏),一种简单有效的自蒸馏方法,无需外部模型或真实答案即可改进VLM。关键在于干净与损坏输入间的预测差异自然产生自监督信号。在NOPD中,模型从损坏输入学习,同时将干净输入下自身预测用作token级监督。在五个视觉推理任务上展示了NOPD的有效性,能匹配甚至超越强化学习方法或外部模型蒸馏。如用Geometry3K的2.1K样本训练时,NOPD使Qwen2.5-VL-7B在验证集上提高20分,在分布外测试集上也有泛化能力,在MathVista上提高7.4分。还证明NOPD是增强VLM的通用方法,在12个基准测试中对三个模型都有改进。
英文摘要
Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinforcement learning with human feedback, or verifiable answers. This limits their ability to improve without external supervision. To tackle this, we propose NOPD (Noisy Student On-Policy Self-Distillation), a simple yet effective self-distillation approach that improves VLMs without any external models or ground-truth answers. Our key insight is that prediction discrepancies between clean and corrupted inputs naturally induce a self-supervision signal. In NOPD, the model learns from corrupted inputs while using its own predictions under clean inputs as token-level supervision. We show the effectiveness of NOPD on five visual reasoning tasks; it can match and even outperform reinforcement learning approaches or distillation from external models. Notably, when trained with 2.1K samples from Geometry3K, NOPD improves Qwen2.5-VL-7B by 20 points on its validation set. It also shows generalization on out-of-distribution test sets and achieves 7.4 point gains on MathVista. Furthermore, we demonstrate that NOPD is a general approach to enhance VLMs, achieving improvements across three models on 12 benchmarks.