arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28590cs.CVcs.CL

VAD:为多模态在线策略蒸馏中的目标重建归因视觉证据

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

发表机构上海交通大学 · 小红书公司 · 香港中文大学
另 2 家 · 查看机构详情
  • Shanghai Jiao Tong University(上海交通大学)
  • Xiaohongshu Inc.(小红书公司)
  • The Chinese University of Hong Kong(香港中文大学)
  • Zhejiang University(浙江大学)
  • Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出视觉归因蒸馏(VAD)算法,通过反事实目标重建分离教师校正中的视觉证据分量,在6个4B/9B规模的细粒度视觉基准上,性能优于现有蒸馏方法。

中文摘要 AI 辅助

多模态在线策略蒸馏(Online Policy Distillation, OPD)通过特权视角教师对学生生成的轨迹进行监督,以传递细粒度视觉知识。然而,其下一个词的校正结果是源混合的,结合了视觉信号、语言先验以及教师特有的效应。关键挑战在于,需估计哪些校正由视觉证据支撑,而非仅确定蒸馏的位置或强度。我们提出视觉归因蒸馏(Visual Attribution Distillation, VAD),这是一种反事实目标重建算法,用于估计教师校正中可归因于视觉的部分。在每个学生生成的前缀处,VAD通过保留和移除相关证据来评估同一个固定教师,中心对数概率的相应变化定义了u_t,这是视觉证据方向的有符号代理,用于估计证据对候选词的揭示程度(支持或反驳)。VAD将原始校正投影到该代理上,得到与干预对齐的分量和代理无法解释的残差,随后从前者重建以学生为锚定的目标。训练期间,该重建目标提供主要监督信号,而特权教师则贡献弱正则化项。在6个细粒度视觉基准(4B和9B规模)上,VAD的性能优于直接特权视角蒸馏和视觉优势加权。词级和受控目标分析表明,代理对齐分量富含任务相关的视觉校正,并产生更强的目标偏移,尤其是当证据反驳错误答案时。这些结果表明,反事实目标重建是源混合监督的有效替代方案。

英文摘要

Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.

补充信息

↑