arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoPID:视觉-语言模型中视觉信息的分解与引导

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

Seulgi Kim, Zhixiong Zhang, Xinwei Zhang, Jie Ling, Ronn Shaw

arXiv 2610.08401首次发表:更新:

发表机构

Georgia Institute of Technology; Amazon(佐治亚理工学院; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无需训练的GeoPID框架,从几何角度分解VLM中的多模态信息,并通过增强视觉特有子空间提升视觉基础能力,在22个模型和14个基准上平均相对准确率提升7.63%。

AI 中文摘要

尽管近期视觉-语言模型(VLMs)在多种应用中展现出卓越性能,但它们往往倾向于低估视觉信息的作用,而过度依赖文本上下文。在本工作中,我们提出了GeoPID,一个无需训练的框架,从几何视角分析VLM中的多模态信息。GeoPID通过视觉与文本表示子空间之间的几何关系,将信息分解为冗余、模态特有和协同三个组成部分。通过对22个VLM和14个基准的广泛分析,我们确认:当问题强烈需要视觉基础时,正确的预测表现出更强的视觉特有成分。基于这一几何分析,我们引入了一种有针对性的干预技术,在推理过程中选择性地增强视觉特有子空间上的视觉表示。由此,在不进行任何额外模型参数更新的情况下,视觉基础能力得到增强,平均相对准确率提升了7.63%。

英文摘要

While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑