arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26093cs.CV

RECAP:用于检测视觉语言模型中空间关系幻觉的关系证据校准

RECAP: Relation Evidence Calibration for Detecting Spatial Relation Hallucinations in Vision-Language Models

Feixiang Liu, Qiang Qiu, Qingyang Li, Hui Xu

首次发表
浏览论文内容

中文总结 AI 辅助

RECAP通过校准关系证据来检测视觉语言模型中的空间关系幻觉,在多个基准上显著降低错误接受率并提升准确率。

中文摘要 AI 辅助

视觉语言模型即使图像支持不兼容的关系,也能自信地回答空间关系问题。我们提出了关系接地的选择性预测:通过审计视觉支持来接受或拒绝已生成的“是/否”答案,而不是将不确定性视为证据。我们的关系证据校准框架RECAP,比较一个主张、其语义矛盾以及可选单侧支持的图像条件似然,然后将这些证据转化为基于答案的拒绝风险。一个仅校准的门控在置信度可证明有信息时将其作为否决权保留,否则仅部署关系证据。在20个组/图像不相交的分割中,RECAP在VSR上将H-FPR@80相对于置信度降低了2.0至17.9个百分点,并在What'sUp上将Qwen3-VL-8B、InternVL3.5-8B和LLaVA-1.5-7B的Acc@80分别提高了3.0、8.6和12.6个百分点。在所有六种设置中,它在所有四个主要指标上均优于匹配的VCD风格视觉对比。全池VSR回退、目标排序的GSR-Bench迁移、等预算监督对照以及两个额外的检查点显示了一致的操作原则:当置信度错位时,结构化反证据补充置信度,而当置信度已有用时,门控保留置信度。

英文摘要

Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-grounded selective prediction: accept or reject an already-produced yes/no answer by auditing its visual support, rather than treating uncertainty as evidence. RECAP, our relation-evidence calibration framework, compares image-conditioned likelihoods for a claim, its semantic contradictions, and optional one-sided supports, then converts these witnesses into an answer-conditioned rejection risk. A calibration-only gate preserves confidence as a veto when confidence is demonstrably informative and otherwise deploys relation evidence alone. Across 20 group/image-disjoint splits, RECAP lowers H-FPR@80 over confidence by between 2.0 and 17.9 points on VSR and raises Acc@80 by 3.0, 8.6, and 12.6 points on What'sUp for Qwen3-VL-8B, InternVL3.5-8B, and LLaVA-1.5-7B. It outperforms matched VCD-style visual contrast on all four primary metrics in all six settings. Full-pool VSR fallback, target-ranked GSR-Bench transfer, equal-budget supervised controls, and two additional checkpoints show a consistent operating principle: structured counterevidence complements certainty when confidence is misaligned, while the gate retains confidence when it is already useful.

补充信息

↑