arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MissingBench-Verified:探究视觉语言模型检测缺失物体部分的无能

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou, Guoping Luo, Shan Du

arXiv 2607.18673首次发表:更新:

发表机构

University of British Columbia; Weathon Software(英属哥伦比亚大学; 威盛软件)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型检测缺失物体部分的能力,提出MissingBench-Verified基准,发现十个领先模型在此场景存在高失败率,现有缓解策略效果不佳,揭示当前VLM在检查监测任务中的根本局限,强调架构或训练干预的必要。

AI 中文摘要

视觉语言模型(VLM)因在图像中虚构不存在的物体而闻名。对于有缺失部分的物体,VLM面临独特挑战,源于现实世界知识偏差和训练数据中此类图像的稀缺。我们提出MissingBench-Verified基准,用于评估特定且实际相关的场景:当视觉语言模型无法识别物体的关键部分被移除时。在十个领先模型中,我们观察到即使外部工具证据与模型视觉感知明显矛盾,失败率仍持续且显著。我们还研究了给予模型图像处理工具(如裁剪、对比度调整)是否能自主检查解决这些失败。结果发现现有缓解策略改善甚微,表明当前VLM在检查和监测任务中有根本局限,凸显架构或训练层面干预的必要性。

英文摘要

Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such images in training data. We present MissingBench-Verified, a benchmark designed to evaluate a specific and practically relevant scenario: when vision-language models fail to recognize that an essential component of an object has been removed. Across ten leading models, we observe consistent and significant failure rates that persist even when external tool evidence explicitly contradicts the model's visual perception. We further ask whether granting models access to image processing tools (e.g., cropping, contrast adjustment) enables autonomous inspection to resolve these failures. We find that existing mitigation strategies, including tool-assisted verification, autonomous visual reasoning, longer reasoning durations, and fine-tuning on an easier dataset, provide negligible improvement, indicating that this failure mode cannot be addressed through current prompting or post-hoc correction techniques. Our findings highlight a fundamental limitation of current VLM for inspection and monitoring tasks and underscore the need for architectural or training-level interventions that enable models to override internal expectations when confronted with contradictory evidence.

CommentsSubmitted to the ECCV 2026 Workshop on Explainable Computer Vision (eXCV). 11 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑