多模态大语言模型中的不确定性感知决策
Uncertainty-Aware Decision Making in Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
本调查围绕决策中心框架,综述多模态大语言模型(MLLMs)的不确定性感知决策相关研究,对比同类调查并指出源感知分解等开放问题,核心是不确定性需改善多模态证据下的系统行为。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models, MLLMs)越来越多地回答其正确性依赖于视觉、文本、时间、声学、文档、图表或具身证据的问题,因此它们的失败不仅是语言层面的。流畅的回答可能掩盖输入质量差、感知错误、基础薄弱、模态间冲突、推理不稳定、分布偏移或无法从提供的证据中回答的问题。本调查围绕以决策为中心的框架组织了关于不确定性感知MLLMs的文献:不确定性来源产生可观测信号,必须对信号进行校准或控制以应对风险,校准后的不确定性应决定系统行动。我们综述了关于token与logit不确定性、语义分歧、扰动不稳定性、基础与归因分数、口头置信度、验证器与评判器分数、保形预测、选择性回答、弃权(不执行)、澄清、检索、自我检查及升级的研究。核心论点是,不确定性不应仅作为置信度数值进行评估,而应评估其是否在证据不足、冲突、偏移或高风险的多模态证据下改善了行为。我们将本调查与仅文本的不确定性和弃权(不执行)调查、广泛的MLLM调查、MLLM幻觉调查以及面向安全的综述进行对比,最后总结了源感知分解、行动感知基准、偏移下的校准、黑盒不确定性估计、更广泛的模态覆盖、可复现报告及以人为中心的不确定性通信等开放问题。
英文摘要
Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a question that is not answerable from the supplied evidence. This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action. We review work on token and logit uncertainty, semantic disagreement, perturbation instability, grounding and attribution scores, verbalized confidence, verifier and judge scores, conformal prediction, selective answering, abstention, clarification, retrieval, self-checking, and escalation. The central argument is that uncertainty should not be evaluated only as a confidence number; it should be evaluated by whether it improves behavior under insufficient, conflicting, shifted, or high-risk multimodal evidence. We position this survey against text-only uncertainty and abstention surveys, broad MLLM surveys, MLLM hallucination surveys, and safety-oriented reviews. We conclude with open problems in source-aware decomposition, action-aware benchmarks, calibration under shift, black-box uncertainty estimation, broader modality coverage, reproducible reporting, and human-centered uncertainty communication.
发表机构
- Khalifa University of Science and Technology(哈利法科技大学)
机构由 AI 辅助整理,请以论文原文为准。