KVE-KD:面向视觉语言模型的关键视觉证据引导知识蒸馏
KVE-KD: Key Visual Evidence-Guided Knowledge Distillation for Vision-Language Models
浏览论文内容
中文总结 AI 辅助
KVE-KD提出动态聚焦任务相关视觉标记的知识蒸馏框架,通过语义锚点和归一化熵选择关键视觉证据,在六个基准上超越现有方法且无推理开销。
中文摘要 AI 辅助
知识蒸馏对于在资源受限设备上部署视觉语言模型至关重要。然而,现有方法通常对视觉标记施加统一的监督,或依赖静态标记选择,这会将任务相关线索与背景噪声混淆,并削弱跨模态推理能力。为解决这一局限,我们提出了关键视觉证据引导的知识蒸馏(KVE-KD),该框架动态地将特征蒸馏聚焦于由教师模型识别的任务相关视觉标记上。具体而言,KVE-KD将预生成阶段的最终文本标记指定为统一语义锚点,并通过迭代移除视觉标记贡献来分析锚点表示的变化,从而确定目标跨模态融合层。在该层内,KVE-KD通过锚点条件注意力分布对视觉标记进行排序,并利用归一化熵选择信息量最大的视觉标记作为关键视觉证据。关键视觉证据随后引导聚焦的视觉特征蒸馏,使学生模型与教师模型的任务相关视觉表示紧密对齐,同时抑制无关背景信息。在六个基准上的大量实验表明,KVE-KD优于最先进的跨模态蒸馏方法,在需要复杂推理和细粒度视觉理解的任务上尤其显著。重要的是,这些改进是在不引入任何推理时开销的情况下实现的。源代码可在该https URL获取。
英文摘要
Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues with background noise and degrades cross-modal reasoning. To address this limitation, we propose Key Visual Evidence-guided Knowledge Distillation (KVE-KD), a framework that dynamically focuses feature distillation on task-relevant visual tokens identified by the teacher model. Specifically, KVE-KD appoints the final pre-generation textual token as a unified semantic anchor and identifies the target cross-modal fusion layer by analyzing changes in the anchor representation through iterative visual-token contribution removal. Within this layer, KVE-KD ranks visual tokens via the anchor-conditioned attention distribution and selects the most informative visual tokens as key visual evidence with normalized entropy. The key visual evidence subsequently guides focused visual feature distillation, making the student align closely with the teacher's task-relevant visual representations while suppressing irrelevant background information. Extensive experiments on six benchmarks demonstrate that KVE-KD outperforms state-of-the-art cross-modal distillation methods, with particularly pronounced gains on tasks requiring complex reasoning and fine-grained visual understanding. Importantly, these improvements are achieved without introducing any inference-time overhead. The source code is available at https://github.com/zhangjianbin07/KVE-KD.
发表机构
- City University of Macau(澳门城市大学)
- Tongji University(同济大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。