MM-OPD:迈向感知与推理之间的又一个瓶颈
MM-OPD: Towards One More Bottleneck Between Perception and Reasoning
浏览论文内容
中文总结 AI 辅助
本文发现图像输入与符号视图输入之间存在“符号视觉差距”,提出MM-OPD在线自蒸馏框架,通过符号到视觉校正引导模型选择正确视觉证据,提升多模态推理能力。
中文摘要 AI 辅助
近期多模态大语言模型(MLLMs)通过同时增强感知与推理能力来推进视觉推理,隐含地假设了一个从感知到推理无缝过渡的过程。然而,我们观察到一个违反直觉的现象,对该假设提出了挑战:在保持模型、问题和解码固定的情况下,我们将图像替换为其标题或代码表示(符号视图),鉴于清晰的图像结构,这似乎是冗余的,但性能却出人意料地提升了10.2%至23.6%,跨越不同模型规模和数据集。我们将这一性能差距称为“符号视觉差距”,并对其进行了更深入的研究。通过实验,我们发现,尽管对于图像输入模型,视觉证据已经可以出现在推理轨迹中,但符号视图输入模型对正确证据的关注度远高于图像输入模型。这表明,尽管当前工作在感知和推理本身方面具备良好能力,但在感知与推理之间,选择感知到的视觉信息作为后续推理的适当证据方面,存在另一个瓶颈。为了解决这一瓶颈,由于符号视图能将注意力引向正确证据,且易于大规模获取,它为证据选择提供了监督,而无需人工标注证据。基于此,我们引入了MM-OPD,一种用于符号到视觉校正的多模态在线自蒸馏框架,通过残差令牌级目标将来自符号条件行为的指导转移到图像条件策略,引导模型关注正确的视觉证据。跨基准和模型规模的实验表明,MM-OPD提升了广泛的多模态能力,在视觉感知、图表与文档理解、数学推理和通用VQA方面均有增益。
英文摘要
Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question, and decoding fixed, we replace images with their caption or code representations (symbolic views), which seems to be redundant given the clear image structures, but the performance surprisingly improves by 10.2% to 23.6% across model scales and datasets. We term this performance gap as the Symbolic Visual Gap and then take a closer look at it. Through experiments, we find that although the visual evidence can already appear in the reasoning trace for the image-input model, the symbolic-view-input model shows much higher attention to the correct evidence than the image-input model. This suggests that despite good capabilities from current works in perception and reasoning themselves, another bottleneck exists between perception and reasoning in selecting perceived visual information as appropriate evidence for subsequent reasoning. To handle this bottleneck, since the symbolic view steers attention toward correct evidence and is readily obtained at scale, it provides supervision for evidence selection without manually labeled evidence. Building on this, we introduce MM-OPD, a multimodal on-policy self-distillation framework for symbolic-to-visual correction that transfers guidance from symbolic-conditioned behavior to the image-conditioned policy through residual token-level targets, steering the model toward correct visual evidence. Experiments across benchmarks and model scales show that MM-OPD improves a broad range of multimodal abilities, with gains in visual perception, chart and document understanding, mathematical reasoning, and general VQA.