在高分辨率多模态大语言模型中更多关注文本
Pay More Attention To Text In High-Resolution MLLMs
浏览论文内容
中文总结 AI 辅助
针对高分辨率多模态大语言模型中问题文本未指定视觉证据的瓶颈,提出无需训练的EviSpec编译器,通过补充证据规格提升定位性能,在多个基准上取得显著相对提升。
中文摘要 AI 辅助
高分辨率多模态大语言模型(MLLMs)的失败通常被归因于视觉问题,这促使研究者采用缩放、裁剪及相关视觉干预手段来恢复细粒度证据或抑制干扰。然而,近期研究表明,相关的视觉证据已编码在中间表示中,这表明仅靠视觉侧的改进是不够的。这引出一个自然的问题:剩余的瓶颈是否在于引导视觉搜索的文本?我们识别出一个此前被忽视的语言瓶颈:为回答问题而构建的问题表述,并不必然指定定位所需的视觉证据。为解决这一不匹配,我们引入了EviSpec,一种无需训练的编译器,它在保留原始问题用于最终推理的同时,推导出互补的证据规格。我们进一步通过匹配对照实验来验证其有效性,这些实验隔离了证据规格与定位的作用。在搜索预算固定的情况下,结构化证据规格相比通用请求取得了8.6%的相对提升。在证据几何匹配的情况下,由EviSpec定位的证据相比随机证据取得了14.8%的相对提升。这些对照实验共同隔离了指定要寻找何种证据而非仅仅扩大视觉访问的益处。在所有五个多模态大语言模型上,EviSpec在三个基准的每一个上都持续优于相应的基线,在V*Bench、HR-Bench-4K和HR-Bench-8K上分别取得了10.4%、8.8%和12.4%的平均相对提升。除了高分辨率推理,EviSpec还在VQA和幻觉聚焦基准上达到了最先进的性能。
英文摘要
Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.
发表机构
- Sichuan University(四川大学)
- Peking University(北京大学)
- Southwest University of Finance and Economics(西南财经大学)
机构由 AI 辅助整理,请以论文原文为准。