arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27890cs.CVcs.AI

RelCheck:用于VLM幻觉校正的双证据空间基础

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

Siddhi Patil, Navrati Saxena, William B. Andreopoulos

首次发表
浏览论文内容

中文总结 AI 辅助

RelCheck提出一种无需训练的事后校正流程,结合场景图三元组和空间谓词双证据,提升VLM空间幻觉校正,在MME位置子任务上显著优于基线。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)经常生成与输入图像不一致的文本。虽然对象级和属性级幻觉已受到相当关注,但关系幻觉(对对象间空间或交互关系的错误描述)在很大程度上仍未得到现有事后校正方法的解决。我们提出了RelCheck,一种无需训练的事后校正流程,通过双关系证据增强对象级视觉基础:来自RelTR的学习场景图三元组和来自边界框几何的确定性空间谓词。这些与Woodpecker风格的对象声明层结合,形成三层视觉知识库,语言模型校正器利用该知识库重写幻觉文本。在LLaVA v1 13B上评估,RelCheck的总MME幻觉得分为630.0,而Woodpecker风格基线为585.0,在位置子任务上增益最大(+31.7分,准确率+从0.367提升至0.600)。四项配置消融实验证实两个关系层独立贡献(McNemar p = 0.025)。这些结果表明,结构化关系证据显著改善了当前MLLMs最欠缺的空间推理子任务上的事后幻觉校正。

英文摘要

Multimodal large language models (MLLMs) fre- quently generate text that is inconsistent with the input image. While object- and attribute-level hallucinations have received considerable attention, relational hallucinations (incorrect de- scriptions of spatial or interactive relationships between objects) remain largely unaddressed by existing post-hoc correction methods. We present RelCheck, a training-free post-hoc correction pipeline that augments object-level visual grounding with dual relational evidence: learned scene-graph triples from RelTR and deterministic spatial predicates from bounding-box geometry. These combine with a Woodpecker-style object claim layer to form a three-layer visual knowledge base, which a language model corrector uses to rewrite hallucinated text. Evaluated on LLaVA v1 13B, RelCheck achieves a total MME hallucination score of 630.0 versus 585.0 for a Woodpecker-style baseline, with the largest gain on the position subtask (+31.7 points, accuracy+ improving from 0.367 to 0.600). A four-configuration ablation confirms that both relational layers contribute independently (McNemar p = 0.025). These results show that structured relational evidence meaningfully improves post-hoc hallucination correction on the spatial reasoning subtasks where current MLLMs are most deficient.

发表机构

  • San José State University(圣何塞州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑