arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关注推理中的关键要素:通过选择性概率质量集中对齐跨模态注意力

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu

arXiv 2609.29940首次发表:更新:

发表机构

University of Technology Sydney; East China Normal University; The University of New South Wales; Shanghai Jiaotong University; The Education University of Hong Kong(悉尼科技大学; 华东师范大学; 新南威尔士大学; 上海交通大学; 香港教育大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出选择性概率质量集中(sPMC)训练框架,通过识别并正则化视觉接地响应注意力头,引导文本到图像注意力集中于语义相关区域,在6个基准上平均零样本提升3%,最高达11.3%。

AI 中文摘要

多模态大语言模型(MLLMs)在视觉推理任务上取得了强劲性能,但仍容易产生幻觉并过度依赖语言先验,常常在未充分使用任务相关视觉证据的情况下生成答案。现有方法主要通过面向推理的监督或推理时策略来改进推理。在本工作中,我们研究一个互补性问题:能否在不直接监督推理过程的情况下,通过增强隐式视觉接地来改进多模态推理?受注意力头功能专门化的启发,我们探究是否可以通过仅引导那些对视觉证据接地最敏感的注意力头来改进推理。我们提出选择性概率质量集中(sPMC),一个训练框架,它识别接地响应头并选择性地正则化其文本到图像的注意力。sPMC将视觉标记上的归一化注意力视为空间概率分布,并利用分割导出的空间先验鼓励概率质量分配给语义相关区域。自适应头选择将这种引导限制在视觉响应头上,同时保持其余头不受约束以保留其互补功能。在6个多模态基准套件上,sPMC在多个MLLMs上实现了平均零样本改进3%,提升最高达11.3%,同时仅正则化其3%-15%的注意力头。这些结果表明,对稀疏和隐式视觉证据路径的定向引导可以直接改进多模态推理。

英文摘要

Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑