arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29193cs.LG

HalluPrism:多模态不确定性应用于诊断而非决策

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

  • Indian Institute of Technology Delhi(德里印度理工学院)
  • Dhirubhai Ambani University(德鲁巴伊·安巴尼大学)

机构由 AI 辅助整理,请以论文原文为准。

Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty

中文总结 AI 辅助

该研究提出HalluPrism,通过视觉扰动敏感性等三维特征诊断多模态大语言模型的失败类别,提升了失败检测的AUROC,明确多模态不确定性应先诊断再决策。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)会对因不同原因失败的答案赋予相似的置信度。我们提出HalluPrism,这是一种行为诊断方法,它在视觉退化、空白图像替换以及接地或关系检查后重新运行答案。这些针对性探测会产生关于视觉扰动敏感性(V)、图像移除置信度保留率(L)和接地/关系探测不稳定性(A)的特征。在来自四个基准测试和四个MLLMs的58000多个示例中,图像移除置信度保留最为普遍,而接地/关系探测不稳定性能更好地区分失败类别。48项源-目标检查中仅有18项是对角对齐的,因此应将这些坐标作为一个整体来解释,而非作为独立的因果来源。在数据集固定的情况下,该联合特征使HallusionBench上的失败类别AUROC从0.634提升至0.769,VizWiz上从0.707提升至0.817,在POPE和VSR上提升幅度较小。在合并的XGBoost分析中,使用标量置信度时AUROC为0.78,使用(V、L、A)时升至0.95,加入置信度后升至0.97。相同的特征并不能自动提升正确性排名,测试的三种直接标量化方法反而会损害其性能。这些结果将失败诊断与弃权(不执行)评分区分开来:多模态不确定性应先表征失败结构,再用于决定是否弃权或修正。

英文摘要

Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confidence retention (L), and grounding/relation-probe instability (A). Across 58K+ examples from four benchmarks and four MLLMs, image-removal confidence retention is most prevalent, while grounding/relation-probe instability better separates failure families. Only 18 of 48 source-target checks are diagonally aligned, so the coordinates should be interpreted jointly rather than as independent causal sources. With the dataset fixed, the joint signature improves failure-family AUROC from 0.634 to 0.769 on HallusionBench and from 0.707 to 0.817 on VizWiz, with smaller gains on POPE and VSR. In pooled XGBoost analysis, AUROC rises from 0.78 with scalar confidence to 0.95 with (V, L, A) and 0.97 when confidence is added. The same signature does not automatically improve correctness ranking. The three tested direct scalarizations can harm it. These results separate failure diagnosis from abstention scoring: multimodal uncertainty should characterize failure structure before it is used to decide whether to abstain or correct.

补充信息

↑