arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭示多模态大语言模型中的认知不确定性:基于因果不变掩码的方法

Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

Haoyang Luo, Linwei Tao, Jie Gui, Xinghao Chen, Chang Xu, Jianyuan Guo, Minjing Dong

arXiv 2610.02887首次发表:更新:

发表机构

City University of Hong Kong; Apple; Southeast University; Huawei Noah’s Ark Lab; University of Sydney(香港城市大学; 苹果公司; 东南大学; 华为诺亚方舟实验室; 悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型的幻觉问题,提出因果不变掩码(CIM)和语义散度度量认知不确定性,并引入快速代理指标EED,在多种基准上达到最先进性能且加速近50%。

AI 中文摘要

多模态大语言模型(MLLMs)存在幻觉问题,这迫切需要对不确定性量化(UQ)以确保可靠的部署。然而,现有方法难以检测由表面关联引起的不确定性,尤其是在查询相关信号较弱的情况下。我们主要将这一问题归因于它们偏向于由数据模糊性引起的偶然不确定性,而忽视了源于模型局限性的认知不确定性。为了进一步分解不确定性类型以实现全面的UQ,我们提出了因果不变掩码(CIM),该方法衡量原始预测与基于因果聚焦视角的条件预测之间的语义偏移。基于这一框架,我们引入语义散度作为UQ的核心度量,并提供理论证据表明其收敛于模型对非因果相关性敏感性的方差,从而确立其捕捉MLLM局限性的能力。为了加速MLLM中的UQ,我们进一步提出了期望嵌入漂移(EED),一种快速的几何代理度量,直接在超球面嵌入空间中估计语义偏移。实验表明,我们的方法在各种基准上达到了最先进的性能,而所提出的EED在性能相当的情况下加速了近50%。

英文摘要

Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.

CommentsAccepted by NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑