arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17269cs.CV

语义-空间一致性验证用于缓解多模态大语言模型中的对象幻觉

Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models

  • Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))
  • Shandong Computer Science Center (National Supercomputer Center in Jinan)(山东省计算中心(国家超级计算济南中心))
  • China Telecom Digital Intelligence Technology Co., Ltd.(中国电信数字智能科技有限公司)
  • Shenyang Aerospace University(沈阳航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Ziheng Ren, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao

AI总结:

提出无需训练的语义-空间一致性验证方法,通过跨查询语义稳定性和区域一致性为对象声明提供可解释视觉证据,有效缓解多模态大语言模型的对象幻觉。

AI中文摘要:

多模态大语言模型从视觉输入生成自然语言响应,但可能提及图像中不存在的对象。在用药辅助、无障碍感知和环境决策中,此类幻觉可能造成现实世界的安全风险。我们提出语义-空间一致性验证(SSAV),一种无需训练的对象声明验证方法。视觉上有依据的声明应在语义等价的查询间保持稳定,并反复定位到同一图像区域。SSAV聚合多个提示以估计语义支持并降低对查询措辞的敏感性。查询诱导的区域验证(QIRV)结合跨查询区域持久性、空间重叠和相对候选优势,以识别孤立的高响应和分散的定位。几何平均融合语义和空间证据,当任一分支缺乏支持时降低验证分数。在三个基础模型和多种评估协议上的实验表明,SSAV有效缓解对象幻觉。在LLaVA-1.5-7B上,在POPE Popular和Adversarial下,COCO、A-OKVQA和GQA上的平均准确率分别提升1.81和3.17个百分点,而CHAIRs从49.40%降至32.80%。这些结果表明,跨查询语义稳定性和区域一致性为对象声明提供了可解释的外部视觉证据。

英文摘要:

Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible perception, and environmental decision-making, such hallucinations can create real-world safety risks. We propose Semantic-Spatial Agreement Verification (SSAV), a training-free method for verifying object claims. A visually grounded claim should remain stable across semantically equivalent queries and repeatedly localize to the same image region. SSAV aggregates multiple prompts to estimate semantic support and reduce sensitivity to query wording. Query-Induced Regional Verification (QIRV) combines cross-query region persistence, spatial overlap, and relative candidate dominance to identify isolated high responses and dispersed localizations. A geometric mean fuses semantic and spatial evidence, lowering the verification score when either branch lacks support. Experiments on three base models and multiple evaluation protocols show that SSAV effectively mitigates object hallucination. On LLaVA-1.5-7B, accuracy averaged across COCO, A-OKVQA, and GQA improves by 1.81 and 3.17 percentage points under POPE Popular and Adversarial, respectively, while CHAIRs decreases from 49.40% to 32.80%. These results show that cross-query semantic stability and regional consistency provide interpretable external visual evidence for object claims.

↑