VAD:为多模态在线策略蒸馏中的目标重建归因视觉证据
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
另 2 家 · 查看机构详情
- Shanghai Jiao Tong University(上海交通大学)
- Xiaohongshu Inc.(小红书公司)
- The Chinese University of Hong Kong(香港中文大学)
- Zhejiang University(浙江大学)
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出视觉归因蒸馏(VAD)算法,通过反事实目标重建分离教师校正中的视觉证据分量,在6个4B/9B规模的细粒度视觉基准上,性能优于现有蒸馏方法。
中文摘要 AI 辅助
多模态在线策略蒸馏(Online Policy Distillation, OPD)通过特权视角教师对学生生成的轨迹进行监督,以传递细粒度视觉知识。然而,其下一个词的校正结果是源混合的,结合了视觉信号、语言先验以及教师特有的效应。关键挑战在于,需估计哪些校正由视觉证据支撑,而非仅确定蒸馏的位置或强度。我们提出视觉归因蒸馏(Visual Attribution Distillation, VAD),这是一种反事实目标重建算法,用于估计教师校正中可归因于视觉的部分。在每个学生生成的前缀处,VAD通过保留和移除相关证据来评估同一个固定教师,中心对数概率的相应变化定义了u_t,这是视觉证据方向的有符号代理,用于估计证据对候选词的揭示程度(支持或反驳)。VAD将原始校正投影到该代理上,得到与干预对齐的分量和代理无法解释的残差,随后从前者重建以学生为锚定的目标。训练期间,该重建目标提供主要监督信号,而特权教师则贡献弱正则化项。在6个细粒度视觉基准(4B和9B规模)上,VAD的性能优于直接特权视角蒸馏和视觉优势加权。词级和受控目标分析表明,代理对齐分量富含任务相关的视觉校正,并产生更强的目标偏移,尤其是当证据反驳错误答案时。这些结果表明,反事实目标重建是源混合监督的有效替代方案。
英文摘要
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.