发表机构
Toyota Motor Europe; University of Oxford(丰田汽车欧洲公司; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统研究多模态大语言模型的免训练不确定性估计,将方法分为令牌级、言语化和语义三类,并在多数据集基准测试中发现各方法在不同回答长度上各有优势。
AI 中文摘要
多模态大语言模型(MLLMs)在广泛的多模态任务中取得了显著性能,然而理解和量化其预测不确定性仍未被充分探索,尽管这对于安全关键应用至关重要。在这项工作中,我们对MLLMs的免训练不确定性量化策略进行了系统性研究,将现有方法归类为三个概念家族:令牌级方法,直接在文本输出空间中操作;言语化方法,通过自然语言提示引发不确定性估计或弃权(不执行)信号;以及语义方法,在语义意义空间中测量不确定性。我们在多个数据集、模型家族、生成代际和规模上对这些策略进行了基准测试,并发现没有单一家族占据主导:令牌级熵(在采样温度1.0下)在短答案上胜出,言语化弃权(不执行)在句子长度响应上胜出,而语义方法在长文本生成上胜出。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of multimodal tasks, yet understanding and quantifying their predictive uncertainty remains underexplored despite being central for safety critical applications. In this work, we present a systematic study of training-free uncertainty quantification strategies for MLLMs, categorizing existing approaches into three conceptual families: token-level methods, which operate directly in the text output space; verbalized methods, which elicit uncertainty estimates or abstention signals via natural language prompts; and semantic methods, which measure uncertainty in a semantic meaning space. We benchmark these strategies across multiple datasets, model families, generations, and scales, and find that no single family dominates: token-level entropy (at sampling temperature 1.0) wins on short answers, verbalized abstention on sentence-length responses, and semantic methods on long-form generation.