发表机构
School of Computer Science and Technology, Tongji University; School of Fashion and Textiles, The Hong Kong Polytechnic University; School of Artificial Intelligence, Shanghai Jiao Tong University(同济大学计算机科学与技术学院; 香港理工大学纺织及制衣学院; 上海交通大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型视觉证据不稳定问题,通过剖析其内部多模态注意力焦点的三阶段分布,提出TRACE框架,该框架能自适应控制推理,在多模型和多基准测试中显著提升基于证据的多模态推理能力。
AI 中文摘要
视觉语言模型在多模态推理基准测试中越来越成功,但其视觉证据一旦进入语言堆栈往往变得不稳定,削弱了基于证据的推理。为理解这种脆弱性,我们通过机制视角研究视觉语言模型的内部动态,发现多模态注意力焦点在深度上有稳定的三阶段重新分布。我们将中间阶段作为视觉中继窗口(VRW),并表明其几何形状随任务需求变化,与基于基础的生成有因果关系,能区分无支持的答案和更强的推理轨迹。在此基础上,我们提出TRACE,一个具有轻量级训练模块的任务自适应推理时间控制框架。它在预填充期间重塑中继分配,并在解码期间交接后保留组装的视觉支持。在四个开放权重的视觉语言模型主干和七个基准测试中,TRACE在对基础敏感的设置上有显著提升,平均提高4.33分,最高提高6.6分,同时也改善了推理繁重的任务。这些结果表明,明确控制跨深度的多模态焦点为加强基于证据的多模态推理提供了一种统一且有效的机制。
英文摘要
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the internal dynamics of VLMs through a mechanistic lens and uncover a stable three-stage redistribution of multimodal attention focus across depth: an early question-conditioned organization, a critical middle visual-dominant relay, and a late return to answer formation. We operationalize the middle phase as the Visual Relay Window (VRW), and show that its geometry varies with task demand, is causally tied to grounded generation, and distinguishes unsupported answers from stronger reasoning trajectories. Guided by this internal rhythm, we propose TRACE, a task-adaptive inference-time control framework with lightweight trained modules. It reshapes relay allocation during prefill and preserves assembled visual support after handoff during decoding. Across four open-weight VLM backbones and seven benchmarks, TRACE delivers large gains on grounding-sensitive settings, improving them by 4.33 points on average and by up to 6.6 points, while also improving reasoning-heavy tasks. These results show that explicitly controlling multimodal focus across depth offers a unified and effective mechanism for strengthening evidence-grounded multimodal reasoning.