发表机构
University of Maryland, College Park; University of California, San Diego; Duke University; MBZUAI(马里兰大学帕克分校; 加利福尼亚大学圣地亚哥分校; 杜克大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出视觉对比自蒸馏(VCSD),将图像内容去除转换为策略内自蒸馏信号,通过对比突出相关候选者,锐化教师原始图像分布并蒸馏到学生中,在多个模型上优于匹配的OPSD,且无需外部教师等额外条件。
AI 中文摘要
策略内自蒸馏(OPSD)很有前景,它去除了策略内蒸馏(OPD)所需的外部教师,但仍需要教师和学生之间的不对称信息。现有方法通过特权答案或视觉证据来创造这种不对称。本文提出是否可以消除这两者,产生一种仅由输入条件驱动的更简单的OPSD形式。为此提出视觉对比自蒸馏(VCSD),将图像内容去除转换为策略内自蒸馏信号。在每个学生生成的响应前缀处,指数移动平均(EMA)教师在相同提示和前缀下产生两个下一个token分布,其token-wise对数概率差异突出由实例级视觉内容特别增加可能性的候选者。利用此对比锐化教师在合理支持范围内的原始图像分布,并将结果全分布目标蒸馏到学生中。使用ViRL39K数据集,VCSD在Qwen3-VL和Qwen3.5模型上始终优于匹配的OPSD。例如,在Qwen3-VL上,2B时七基准聚合从62.27%提高到67.04%,4B时从71.30%提高到73.16%,8B时从72.51%提高到76.26%。此外,VCSD不需要外部教师、特权答案、视觉证据信号、推理痕迹或额外推理时间成本。
英文摘要
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from $62.27\% \rightarrow 67.04\%$ at 2B, $71.30\% \rightarrow 73.16\%$ at 4B, and $72.51\% \rightarrow 76.26\%$ at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.
Comments15 pages