发表机构
SANNO University(山王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态上下文学习中视觉上下文时而被利用时而被忽视的问题,提出VIB-ICL框架,通过CMIG量化跨模态信息,推导理论界并经五组基准实验验证,实现准确率提升与演示样本减少。
AI 中文摘要
大型视觉语言模型展现出强大的上下文学习(ICL)能力,但视觉上下文何时以及为何对多模态ICL有帮助仍未被充分理解。实证研究揭示了一种令人困惑的二分现象:模型有时能有效利用视觉演示,却常常完全忽视它们。我们提出VIB-ICL,这是一种基于信息瓶颈原理解决该二分现象的信息论框架。我们引入跨模态信息增益(CMIG),用于量化视觉上下文在文本上下文之外提供的关于目标的额外互信息。我们推导了一个泛化界,表明多模态ICL相对于纯文本ICL的超额风险由CMIG决定,证明当视觉信息非冗余时,多模态ICL可被证明优于纯文本ICL。我们进一步证明,常被视为失败模式的视觉上下文忽视,在视觉信息冗余时是信息瓶颈最优解,由此得出闭式的注意力重新分配原理,规定应如何自适应调整视觉注意力权重。我们将该原理实例化为VIB-ICL算法,该算法通过变分界估计CMIG并动态重新分配注意力。在五个基准上的实验表明,准确率提升最高达4.7%,所需演示样本减少35%,验证了我们的理论预测。
英文摘要
Large vision-language models exhibit strong in-context learning (ICL) capabilities, yet when and why visual context helps multimodal ICL remains poorly understood. Empirical studies show a puzzling dichotomy: models sometimes effectively leverage visual demonstrations, yet often neglect them entirely. We propose VIB-ICL, an information-theoretic framework that resolves this dichotomy through the Information Bottleneck principle. We introduce the Cross-Modal Information Gain (CMIG), which quantifies the additional mutual information that visual context provides about the target beyond textual context. We derive a generalization bound showing that multimodal ICL's excess risk over text-only ICL is governed by the CMIG, proving that multimodal ICL provably outperforms text-only ICL when visual information is non-redundant. We further prove that visual context neglect, often viewed as a failure mode, is the Information Bottleneck-optimal solution when visual information is redundant, yielding a closed-form Attention Reallocation Principle that prescribes how visual attention weights should be adaptively adjusted. We instantiate this principle in the VIB-ICL algorithm, which estimates CMIG via variational bounds and dynamically reallocates attention. Experiments on five benchmarks demonstrate consistent improvements of up to 4.7\% accuracy gains and 35\% reduction in required demonstrations, validating our theoretical predictions.