arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当有帮助的文本有害时:视觉语言模型中的选项重定向偏差

When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models

Tam Le Thi Thanh, Hoang Tran Van, Hong-Hanh Nguyen-Le, Thanh Duc Ngo

arXiv 2609.32489首次发表:更新:

AI 中文总结

本研究揭示视觉语言模型中,与问题一致但违背图像并支持干扰项的辅助文本会系统性重定向模型预测,导致高达53.1%的准确率下降,并提出无需训练的推理时干预方法以缓解此问题。

AI 中文摘要

在三模态视觉问答(VQA)中,辅助文本通常用于补充视觉和文本输入,但其可靠性往往不受控制。虽然先前的工作研究了普遍的模态冲突,但在固定的图像-问题-选项背景下,不同类型的不可靠辅助文本如何影响答案选择仍不清楚。在这项工作中,我们表明最有害的辅助文本不一定是最事实错误的,而是那些与问题一致但违背图像并偏向特定干扰项的文本,导致模型预测的系统性重定向。为了隔离这一效应,我们引入了文本可靠性阶梯(Textual Reliability Ladder),这是一种受控的诊断协议,将辅助文本沿三个轴分解:图像一致性、问题相关性和选项支持。在多个数据集(ScienceQA、VCR、A-OKVQA、Causal-VidQA)和最近的视觉语言模型(VLMs)上,我们发现这种支持干扰项的文本导致最大的准确率下降(高达53.1%),并将错误集中在特定的错误选项上。为了缓解这种失败模式,我们提出了一种无需训练的推理时干预方法,通过噪声稳定性引导和动态接地(dynamic grounding)显式抵消这种重定向效应,在忠实文本下大幅保持性能的同时减少重定向错误。我们的结果强调,辅助文本的可靠性必须在决策层面理解,而非仅通过事实正确性,并为更稳健的三模态推理提供了实用途径。

英文摘要

In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixed image-question-option contexts. In this work, we show that the most harmful auxiliary text is not necessarily the most factually incorrect, but the one that aligns with the question while contradicting the image and favoring a specific distractor, leading to systematic redirection of model predictions. To isolate this effect, we introduce the Textual Reliability Ladder, a controlled diagnostic protocol that decomposes auxiliary text along three axes: image consistency, question relevance, and option support. Across multiple datasets (ScienceQA, VCR, A-OKVQA, Causal-VidQA) and recent VLMs, we find that such distractor-supporting text induces the largest accuracy drops (up to 53.1%) and concentrates errors on specific incorrect options. To mitigate this failure mode, we propose a training-free inference-time intervention that explicitly counteracts this redirection effect via noise-stability steering and dynamic grounding, reducing redirected errors while largely preserving performance under faithful text. Our results highlight that auxiliary-text reliability must be understood at the decision level, rather than solely through factual correctness, and provide a practical pathway toward more robust tri-modal reasoning.

CommentsAccepted at ACM Multimedia 2026 (ACM MM 2026). 25 pages, 16 figures. This arXiv version includes supplementary material

DOI:10.1145/3767308.3836232

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑