监督之前的感知:来自反事实盲区的自包含视觉蒸馏
Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
查看机构详情
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出首个完全自包含的视觉自蒸馏框架 CVPD,通过识别模型视觉盲区生成密集对比监督,在 Qwen3-VL-8B-Instruct 上的 12 个基准中优于 6 个自进化基线,取得多项指标提升且无性能倒退。
中文摘要 AI 辅助
多模态大语言模型(MLLM)的自改进通常由仅提供粗略标量反馈的基于奖励的方法驱动。蒸馏提供了更丰富的替代方案,通过密集的 token 级监督实现,但在视觉领域,它通常依赖于使用外部注释、工具或更强模型构建的特权上下文。我们引入 CVPD(Contrastive Counterfactual Visual Process Distillation,对比反事实视觉过程蒸馏),据我们所知,这是首个用于 MLLM 的完全自包含的密集、在线策略、token 级视觉自蒸馏框架。CVPD 识别视觉盲区:放大某一区域会改变并锐化模型的答案分布,而移除同一区域则会使全图行为基本不变。这些区域揭示了模型可编码但在全图条件下无法一致利用的感知信息。我们提出三闸门反事实准则,直接从模型自身响应中识别这些区域,并将其转换为用于自蒸馏的密集对比监督。在 Qwen3-VL-8B-Instruct 上,CVPD 在 12 个基准测试中优于 6 个自进化基线,包括依赖外部 GPT-4o 监督的方法,且无任何性能倒退;它在 OCRBench 上提升了 +3.60,在 MMStar 细粒度感知上提升了 +3.38,在 MMStar 逻辑推理上提升了 +3.08,同时在更广泛的多模态基准上保持或提升了性能。
英文摘要
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of $+3.60$ on OCRBench, $+3.38$ on MMStar Fine-Grained Perception, and $+3.08$ on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.