发表机构
The University of Queensland; Shenzhen University(昆士兰大学; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全模态LLMs中模型过度依赖图像回答音频问题的跨模态捷径,提出因子化模态诊断方法定位问题,并设计DMC-Repair训练策略,将图像引起的答案效应降低59.9%,且不损害音频问答性能。
AI 中文摘要
全模态大语言模型(LLMs)被期望使用问题明确指出的模态来回答问题。然而,现有的训练范式很少验证模型是否真正遵循该模态,因为来自同一样本的多模态输入通常为同一答案提供冗余证据。在这项工作中,我们揭示了全模态LLMs中一种普遍存在的跨模态捷径:当被问及与音频相关的问题时,模型对图像的依赖程度与对音频的依赖程度相当,有时甚至更高。为了系统性地诊断这一行为,我们引入了因子化模态诊断方法,该方法独立地在样本之间交换音频和图像,以隔离每种模态的因果贡献。在不同设置下的两个模型家族中,我们发现这种捷径在监督微调和强化学习后训练中持续存在,而基于评判器的强化学习可能进一步放大对无关视觉信息的依赖。基于这一发现,我们提出了DMC-Repair,该方法在同类跨模态交换样本上训练模型,同时根据问题指定的模态分配监督信号。这防止了模型利用同一片段内模态之间的虚假对应关系。实验表明,DMC-Repair将图像引起的答案效应份额降低了59.9%,有效抑制了跨模态捷径,同时不损害音频问题回答性能。捷径依赖性的降低在两个模型家族中普遍适用,并零样本泛化到未见数据集和未见基准,且在后续后训练中持续存在。代码可在以下网址获取:https URL。
英文摘要
Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at https://anonymous.4open.science/r/DMC-Repair.
Comments25 pages, 11 figures, 16 tables