arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

层、汇点与缩放:多模态大语言模型的适应性证据选择

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang, Yan Wang, Zhenwei Zhang

arXiv 2609.16795首次发表:更新:

发表机构

Sichuan University(四川大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型忽视相关证据的问题,提出无需训练的AREA方法,通过探针标记动态决策是否干预、证据量和文本刷新时机,在多个基准上达到最优性能。

AI 中文摘要

多模态大语言模型(MLLMs)通过将图像中的视觉证据与从外部来源检索的事实相结合,能够回答知识密集型的视觉问题。然而,MLLMs 可能会忽视两种模态中的相关证据,对正确答案所需的文本句子或视觉区域的关注较弱。近期的研究通过在生成前突出显示检索到的文本和标记视觉区域来解决这一问题,但采用固定的单次策略,无法适应三个变化来源:突出显示是否必要、不同示例需要多少证据,以及随着答案的展开,不同的文本证据何时变得相关。我们引入了适应性相关性引导的证据分配(AREA),这是一种无需训练、推理时的方法,将证据突出显示表述为适应性分配。AREA 生成一个单一的探针标记,从固定的骨干层读取视觉和文本相关性,然后做出三个决策:i) 是否干预(由自然注意力覆盖和视觉汇点污染控制),ii) 暴露多少证据(由相关性熵决定),以及 iii) 在生成过程中何时刷新文本(由因果上下文注意力峰值触发)。在四个 KB-VQA 和七个标准多模态基准测试中,使用九个冻结的 MLLM 检查点,AREA 在无需训练的突出显示方法中建立了最佳性能。

英文摘要

Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the correct answer. Recent efforts address this by highlighting retrieved text and marking visual regions before generation, but apply a fixed, one-shot policy that cannot adapt to three sources of variation: whether highlighting is necessary, how much evidence different examples require, and when different textual evidence becomes relevant as the answer unfolds. We introduce Adaptive Relevance-guided Evidence Allocation (AREA), a training-free inference-time method that formulates evidence highlighting as adaptive allocation. AREA generates a single probe token to read visual and textual relevance from fixed backbone layers, then makes three decisions: i) whether to intervene (controlled by natural attention coverage and visual sink contamination), ii) how much evidence to expose (determined by relevance entropy), and iii) when to refresh text during generation (triggered by causal context-attention peaks). Across four KB-VQA and seven standard multimodal benchmarks with nine frozen MLLM checkpoints, establishes the best performance among training-free highlighting methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑