AI 中文总结
针对视觉语言模型在线策略蒸馏中教师修正与学生能力不匹配的问题,提出FP-OPD方法,在8B到2B的蒸馏中使7个多模态基准平均得分分别提升2.77和1.60个点。
AI 中文摘要
在线策略蒸馏(OPD)从当前学生策略中采样轨迹,并最小化这些轨迹前缀处学生与教师下一个标记分布之间的标记级差异,使蒸馏状态与学生自身的生成分布对齐。然而,该方法仍假设完整的教师分布是适配所有学生能力的合适目标。在视觉语言推理中,教师的修正可能依赖于紧凑学生无法表示的视觉差异。我们的目标缩放研究表明,当目标趋近于完整教师分布时,学生实现的规定转变更少,下游性能更差。因此,我们提出Fisher投影在线策略蒸馏(FP-OPD),仅蒸馏局部可实现的教师修正。FP-OPD使用连续视觉扰动估计学生的局部视觉切空间,并在学生Fisher度量下将中心化的教师-学生对数概率差距投影到该空间。所得的能力感知目标在学生轨迹上通过全词汇反向KL进行优化,保留了标准OPD框架。在8B到2B的蒸馏中,FP-OPD在所有7个评估的多模态基准上均有提升,将平均得分较预训练学生提高2.77个点,较标准OPD提高1.60个点。这些结果表明,局部可实现的教师修正为蒸馏紧凑视觉语言模型提供了更有效的目标。
英文摘要
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. This aligns the distillation states with the student's own generation distribution. However, it still assumes that the complete teacher distribution is an appropriate target across student capacities. In vision--language reasoning, teacher corrections can depend on visual distinctions that a compact student cannot represent. Our target-scaling study shows that, as the target approaches the complete teacher distribution, the student realizes less of the prescribed shift and obtains worse downstream performance. We therefore propose \emph{Fisher-Projected On-Policy Distillation} (FP-OPD), which distills only locally realizable teacher corrections. FP-OPD uses continuous visual perturbations to estimate the student's local visual tangent space and projects the centered teacher--student log-probability gap onto this space under the student's Fisher metric. The resulting capacity-aware target is optimized with full-vocabulary reverse KL on student trajectories, retaining the standard OPD framework. In 8B-to-2B distillation, FP-OPD improves all seven evaluated multimodal benchmarks. It raises the average score by 2.77 points over the pretrained student and by 1.60 points over standard OPD. These results demonstrate that locally realizable teacher corrections provide a more effective target for distilling compact vision--language models.