发表机构
Harbin Institute of Technology, Shenzhen; Tsinghua Shenzhen International Graduate School, Tsinghua University; Peng Cheng Laboratory; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(哈尔滨工业大学(深圳); 清华大学深圳国际研究生院; 鹏城实验室; 中国科学院深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MLLM推理中感知与逻辑混淆的问题,提出无训练的注意力引导式切换框架,通过视觉-文本注意力比率自适应切换显式与隐式推理,实现性能与效率的双重提升。
AI 中文摘要
多模态大语言模型(MLLM)的推理需要细粒度视觉感知与严谨逻辑推演。基于显式文本的思维链(CoT)计算成本高昂且易产生视觉幻觉,现有隐式推理方法通常需要代价高昂的训练,将无训练的大语言模型(LLM)推理机制直接适配至多模态场景会导致性能不稳定。我们发现该失败源于它们对token级熵的依赖,这从根本上混淆了感知歧义(如模糊的视觉细节)与逻辑不确定性(如复杂推理步骤)。为克服这一瓶颈,我们提出一种面向MLLM的新型无训练推理策略,明确解耦感知与推理。我们提出一种新型指标:视觉-文本注意力比率,用于动态衡量模型的认知焦点。在该指标引导下,我们提出的注意力引导式切换(AGS)框架会自适应地为感知token触发隐式推理,以在连续空间中保留高保真视觉信息,同时为逻辑token强制生成显式文本,以维持结构锚定。大量实验表明,我们的方法达到了SOTA性能,通过减少自回归步骤与延迟,显著提升了准确率与推理效率。代码已发布在此https URL。
英文摘要
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is computationally expensive and prone to visual hallucinations, while existing latent reasoning methods typically require costly training. Furthermore, directly adapting training-free LLM reasoning mechanisms to the multimodal setting yields unstable performance. We identify that this failure stems from their reliance on token-level entropy, which fundamentally conflates perceptual ambiguity (e.g., unclear visual details) with logical uncertainty (e.g., complex reasoning steps). To overcome this bottleneck, we present a novel training-free inference strategy for MLLMs that explicitly decouples perception and reasoning. We propose a novel metric, the vision-to-text attention ratio, to dynamically gauge the model's cognitive focus. Guided by this metric, our proposed framework, Attention-Guided Switching (AGS), adaptively triggers latent reasoning for perceptual tokens to preserve high-fidelity visual information in the continuous space, while enforcing explicit text generation for logical tokens to maintain structural anchoring. Extensive experiments demonstrate that our method achieves state-of-the-art performance, significantly improving both accuracy and inference efficiency by reducing autoregressive steps and latency. Code is released at https://github.com/swordAndSnow/MM26-AGS.
CommentsAccepted by ACM MM 2026. 10 pages, 6 figures, 5 tables