arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当无关文本起作用时:多模态大语言模型中的仿射间隔偏移

When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

Yinfeng Wang, Zhiyuan Yao, Zheren Fu, Lei Zhang, Zhendong Mao

arXiv 2608.19208首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究多模态大语言模型(MLLMs)中任务无关文本对视觉判断任务的影响,发现其会使模型决策间隔发生可预测的仿射变换,为提升模型对噪声上下文的鲁棒性提供了诊断基础。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)常接触辅助文本上下文,其对视觉接地任务的影响尚未得到充分探索。本文在二元视觉判断框架内,通过将任务无关上下文建模为受控干预,研究其影响。在保持提示结构不变的同时改变辅助输入,我们观察到无关文本在各类基准测试中始终会使模型预测产生偏差。为超越性能指标,我们通过二元候选样本间对数概率差定义的决策间隔来表征这种敏感性。分析揭示了一种稳健的几何规律:上下文条件下的间隔遵循其无上下文对应项的一致仿射变换。该发现表明,无关上下文并非表现为非结构化随机噪声,而是模型偏好的可估计畸变。我们进一步将拟合的仿射参数解释为视觉承诺保留和定向答案偏差的度量。这些发现为MLLMs中无关上下文效应提供了间隔层面的诊断视角,并为未来关于噪声上下文鲁棒性的研究提供了基础。

英文摘要

Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑