更好的教师监督是否足够?多模态在线策略蒸馏中学生侧学习的解锁
Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation
- Nanjing University(南京大学)
- Fudan University(复旦大学)
- Nanjing University of Posts and Telecommunications(南京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态在线策略蒸馏中学生感知瓶颈,提出S-OPD框架,通过教师校准策略对比和策略一致性两个目标强化学生视觉学习,无需额外标注或参数,在八个基准上持续提升性能,最高提升4.25分。
AI中文摘要:
在线策略蒸馏(OPD)通过在学生的自身轨迹上提供来自教师的令牌级监督来提升推理能力。现有方法主要侧重于增强这种教师侧指导(例如,通过丰富教师输入和细化教师反馈),然而我们发现,有限的学生感知能力是多模态OPD中的另一个关键瓶颈。通过提供神谕视觉事实,对于弱教师和强教师,OPD训练学生的性能仍能大幅提升。为解决这一瓶颈,我们提出了S-OPD,一个简单的多模态在线策略蒸馏框架,通过两个目标显式强化学生感知学习。具体而言,教师校准策略对比将原始图像和掩码图像下的学生策略进行对比,并采用基于教师的令牌级门控,增强学生在推理过程中对视觉证据的依赖。策略一致性将原始图像和噪声扰动图像下的学生策略进行对齐,进一步提升对视觉噪声的感知鲁棒性。值得注意的是,我们的方法可以无缝插入现有OPD框架,无需额外数据标注、模型参数或推理操作。在跨学生规模和蒸馏范式的八个基准上的大量实验表明,性能持续提升,在LogicVista上最高提升4.25分。当与现有教师侧监督方法结合时,我们的方法能带来进一步增益。代码可在该https URL获取。
英文摘要:
On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at https://github.com/Sirilaw/S-OPD.