将视觉推理蒸馏到文本空间
Distilling Visual Reasoning into Text Space
浏览论文内容
中文总结 AI 辅助
提出V2T框架,通过知识蒸馏将视觉推理内化到文本空间,避免推理时生成中间视觉表示,在多个多模态基准上平均准确率提升14.3%,训练速度提升42倍。
中文摘要 AI 辅助
大型视觉-语言模型(LVLMs)在多模态推理方面展现出巨大潜力,但在处理需要超越输入图像中直接可观察概念的任务时,常常遇到困难。现有方法生成中间图像或潜在视觉标记来引导推理,但这些表示可能引入错误,并随着推理的推进日益干扰文本推理。我们提出视觉到文本思维链蒸馏(V2T),这是一种使LVLMs能够在推理时不生成中间视觉表示而内化视觉推理的框架。V2T首先使用交错的视觉和文本思维链训练一个教师LVLM,然后利用知识蒸馏,通过教师的logits和来自真实文本推理的交叉熵监督来训练学生LVLM。当推理图像可以映射到原始图像时,V2T还可以额外蒸馏教师对相应区域的注意力,而真实边界框可进一步引导后续的强化学习阶段。在多个多模态推理基准上的实验表明,V2T始终优于教师模型和现有基线,在保留集上平均准确率提高14.3%,在更广泛的视觉评估套件上提高2.7%。此外,轻量级SFT和大幅减少的RL使V2T的训练速度比最先进的基线快高达42倍。
英文摘要
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
- Cisco Research(思科研究院)
机构由 AI 辅助整理,请以论文原文为准。