发表机构
Beihang University; University of the Chinese Academy of Sciences(北京航空航天大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对细粒度视觉理解,提出证据对齐的多模态在线自蒸馏(EAD),通过构建证据参考加权教师修正,仅用6%监督质量超越现有方法。
AI 中文摘要
细粒度视觉理解要求模型能够识别复杂图像中的微小细节。多模态在线自蒸馏(OPSD)通过使用以证据为中心的裁剪区域为条件的教师模型,来监督以原始图像为条件的学生模型沿着学生生成的轨迹进行学习,从而解决这一挑战。理想情况下,教师修正,即从学生向特权教师的分布变化,应由任务相关的视觉证据驱动。然而,使教师有效的设计也引入了其他干扰。使用滞后或冻结的教师模型可以提高训练稳定性,但引入了与不断演变的学生之间的模型状态差距,而裁剪增强了任务相关证据但也丢失了视觉上下文。这两种干扰源使得教师修正并非纯粹依赖于视觉证据。我们引入了证据对齐的多模态在线自蒸馏(EAD),它保留以裁剪为条件的教师作为目标,但构建了一个单独的证据参考来加权修正。为了排除滞后模型状态对此参考的影响,EAD使用当前学生来测量预测变化。为了避免裁剪引起的上下文变化,EAD在原始图像中掩蔽证据区域,同时保留其他视觉上下文。学生从掩蔽图像预测到原始图像预测的变化,为视觉证据如何改变学生预测的方向提供了一个受控参考。EAD根据每个教师修正与该参考的余弦对齐度进行加权,即保留对齐的修正并降低其余修正的权重。仅保留密集OPSD监督质量的6%,EAD持续优于先前的最先进方法。
英文摘要
Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories. Ideally, teacher corrections, the distributional changes from the student toward the privileged teacher, should be driven by task-relevant visual evidence. However, the designs that make the teacher effective also introduce other interference. Using a lagged or frozen teacher improves training stability but introduces a model-state gap from the evolving student, while cropping enhances task-relevant evidence but also loses the visual context. These two sources of interference make the teacher corrections not purely rely on the visual evidence. We introduce Evidence-Aligned multimodal on-policy self-Distillation (EAD), which retains the crop-conditioned teacher as the target but constructs a separate evidence reference for weighting the corrections. To exclude the effect of lagged model-state from this reference, EAD measures prediction changes using the current student. To avoid crop-induced context changes, EAD masks the evidence region in the original image while preserving the other visual context. The change from the student's masked-image prediction to its original-image prediction provides a controlled reference for the direction in which the visual evidence shifts the student's prediction. EAD weights each teacher correction by its cosine alignment with the reference, i.e., retaining aligned corrections and downweighting the rest. Retaining only 6\% of the supervision mass of dense OPSD, EAD consistently outperforms previous state-of-the-art methods.