arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪种模态起决定作用?多模态大语言模型的反事实模态归因

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera, David Watson, Senka Krivic

arXiv 2608.00076首次发表:更新:

AI 中文总结

该研究针对多模态大语言模型无法确定预测驱动模态的问题,提出反事实模态归因(CMA)框架,经实验验证其能准确识别决策驱动模态,可用于多模态模型的安全审计。

AI 中文摘要

多模态大语言模型(Multimodal large language models, MLLMs)正通过整合图像与文本的互补信息,越来越多地支持高风险决策。现有可解释性方法虽能识别有影响力的图像区域或文本标记,但无法回答一个根本问题:哪种模态驱动了预测?因此,模型可能在依赖错误证据来源的情况下产生正确输出,掩盖捷径学习与不安全推理。我们将模态归因表述为多模态基础模型的互补可解释性目标,并提出反事实模态归因(Counterfactual Modality Attribution, CMA)——首个量化MLLMs模态层级贡献的框架。CMA利用耦合扩散先验生成仅图像、仅文本及联合多模态反事实,再通过基于夏普利值的合作博弈论公式,将其转化为严谨的模态归因分数。我们在具有真实模态依赖的受控合成基准及真实多模态临床数据集上评估CMA,其在98%的受控案例中能正确识别决策驱动模态,且始终优于基线方法,揭示了仅靠预测准确率无法发现的跨模态推理缺陷。我们的结果确立了模态归因是超越特征归因的可解释性补充维度,为安全关键应用中的多模态基础模型审计提供了严谨框架。

英文摘要

Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction? Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning. We formulate modality attribution as a complementary explainability objective for multimodal foundation models and propose Counterfactual Modality Attribution (CMA), the first framework for quantifying modality-level contributions in MLLMs. CMA generates image-only, text-only, and joint multimodal counterfactuals using coupled diffusion priors and converts them into principled modality attribution scores through a cooperative game-theoretic formulation based on Shapley values. We evaluate CMA on controlled synthetic benchmarks with known ground-truth modality reliance and on a real-world multimodal clinical dataset. CMA correctly identifies the decision-driving modality in 98% of controlled cases and consistently outperforms baselines, revealing failures of cross-modal reasoning that remain invisible to predictive accuracy alone. Our results establish modality attribution as a complementary dimension of explainability beyond feature attribution, providing a principled framework for auditing multimodal foundation models in safety-critical applications.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑