arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38777cs.CV

蒸馏视觉证据,而非仅蒸馏答案:面向视觉语言模型的跨世界在线策略蒸馏

Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models

Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出跨世界在线策略蒸馏(CW-OPD),通过构建视觉证据不同的世界并蒸馏教师信念转变,显式训练学生依赖关键视觉证据,在CWBench上显著提升跨世界一致性。

中文摘要 AI 辅助

视觉语言模型(VLM)蒸馏的一个核心目标是同时迁移教师的语言能力和视觉理解能力。然而,现有方法主要监督学生的输出,使得视觉理解仍处于隐式状态。我们的分析表明,学生可以在不依赖相同视觉证据的情况下匹配教师的答案,这引发了一个问题:我们如何确保学生响应真正决定答案的视觉信息?为此,我们提出了跨世界在线策略蒸馏(CW-OPD),该方法显式监督学生对视觉证据变化的响应。对于每个示例,CW-OPD构建两个共享问题和场景上下文但答案关键证据不同的视觉世界,从而产生不同的答案。我们在两个世界中进行在线策略蒸馏,并蒸馏教师的跨世界信念转变,鼓励学生不仅匹配教师预测的内容,还要匹配其预测随证据变化的原因。梯度分析表明,该术语对两个世界共有的错误具有不变性,并提供了仅端点匹配无法提供的修正信号。通过这种方式,CW-OPD将依赖相关视觉证据作为显式蒸馏目标,而非输出匹配的隐式结果。为了诊断模型是否真正将其答案基于视觉证据,我们引入了CWBench,通过跨世界配对准确率(CWPA)衡量跨世界一致性。在Qwen3.5-4B上的实验表明,CW-OPD平均比最强基线高出1.2个百分点,且4B学生模型在CWBench上的CWPA比DeepSeek-V4.1(552B)高出22.4个百分点。代码已在此https URL发布。

英文摘要

A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match the teacher's answer without relying on the same visual evidence, raising the question: how can we ensure the student responds to the visual information that actually determines the answer? To this end, we propose \textbf{Cross-World On-Policy Distillation (CW-OPD)}, which explicitly supervises the student's response to changes in visual evidence. For each example, CW-OPD constructs two visual worlds that share the question and scene context but differ in answer-critical evidence, yielding different answers. We perform on-policy distillation in both worlds and distill the teacher's cross-world belief transition, encouraging the student to match not only \emph{what} the teacher predicts but also \emph{why} its prediction changes with the evidence. A gradient analysis shows that this term is invariant to errors shared by both worlds and supplies a corrective signal invisible to endpoint matching alone. In this way, CW-OPD makes reliance on the relevant visual evidence an explicit distillation target rather than an implicit consequence of output matching. To diagnose whether a model truly grounds its answers in visual evidence, we introduce CWBench, which measures cross-world consistency via Cross-World Pair Accuracy (CWPA). Experiments on Qwen3.5-4B show that CW-OPD outperforms the strongest baseline by \textbf{1.2} points on average, and the 4B student exceeds DeepSeek-V4.1 (552B) by \textbf{22.4} CWPA points on CWBench. Code is released in https://github.com/baokou-fw2/CWAD.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑