AI 中文总结
提出渐进视角在线蒸馏(PVD),通过填充裁剪逐步将学生视角从区域转向整图,利用区域优势加权迁移教师优势,在多项基准上平均准确率达77.51,显著超越基线。
AI 中文摘要
区域到全局的蒸馏利用裁剪条件引导来提升整图理解能力。其挑战在于如何有效地将教师的基于裁剪的优势迁移到学生的整图推理中。我们提出了渐进视角在线蒸馏(PVD),该方法通过一个保持宽高比的中间填充裁剪,将学生的视角分布从裁剪逐渐转向整图。填充裁剪在保留区域内容的同时,匹配整图的视觉令牌网格。在各个阶段中,视角混合为整图分配递增的概率。一种轻量级的区域优势加权利用基于裁剪条件的师生对数概率差来重新分配令牌级监督。在每次采样输入下进行评估时,当差距较小时施加温和的重新加权,而当差距扩大时则强调差距较大的令牌。一种Jensen-Shannon度量分解将该调度解释为从匹配输入模仿向部署目标的转变。在涵盖感知、视觉数学和通用多模态问答的基准测试中,PVD-full在三个随机种子上达到了77.51的平均准确率,比无奖励蒸馏基线提高了2.01个百分点,比其奖励匹配变体提高了1.00个百分点。在无奖励设置下,PVD-distill仍获得了1.16个百分点的提升。
英文摘要
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.