arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12818cs.CV

打破似曾相识:通过视觉语言推理对视觉场所识别进行独立审计

Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

Sania Waheed, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对视觉场所识别中图像匹配阈值在环境变化下不可靠的问题,引入视觉场所识别审计框架,利用视觉语言模型通过联合推理评估检索匹配,经实验验证该方法有效提升召回率并降低误接受率。

中文摘要 AI 辅助

视觉场所识别(VPR)是机器人应用中精确定位和长期自主导航的关键促成因素,如同时定位与地图构建(SLAM)中的回环检测。然而,现实世界中的VPR部署依赖于选择平衡精度和召回率的图像匹配阈值,这些阈值通常使用标记的验证数据进行调整并在部署期间固定,在无地面真值的环境变化下不可靠。本文引入视觉场所识别审计,这是一个独立的检索后验证框架,利用视觉语言模型(VLM)通过对查询图像和候选图像联合推理来评估检索到的匹配。与传统验证方法不同,该方法无需特定架构的置信度度量、依赖数据集的阈值或部署环境的先验知识即可进行实例级验证。使用五种先进的VPR方法和四种VLM在六个基准数据集上评估了该方法。结果表明,与现有方法相比,基于VLM的审计平均将召回率@1提高了13.6%,同时将误接受率降低到12%,精度保持在95%以上,覆盖率保持在75%以上。

英文摘要

Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-world VPR deployment relies on selecting an image matching threshold that balances precision and recall. These thresholds are typically tuned using labeled validation data and fixed during deployment, making them unreliable under environmental changes where ground truth is unavailable. This is particularly problematic in safety-critical robotics, where accepting a false loop closure can corrupt the estimated trajectory and map. In this work, we introduce Visual Place Recognition Auditing, an independent post-retrieval verification framework that leverages Vision-Language Models (VLMs) to assess retrieved matches by reasoning jointly over query and candidate images. Unlike conventional verification methods, our approach performs instance-level verification without requiring architecture-specific confidence measures, dataset-dependent thresholds, or prior knowledge of the deployment environment. We evaluate our method on six benchmark datasets using five state-of-the-art VPR methods and four VLMs. Results show that VLM-based auditing improves recall@1 by 13.6% on average as compared to state-of-the-art methods while reducing false acceptance rates to 12%, maintaining precision above 95% and coverage above 75%.

发表机构

  • School of Electronics and Computer Science, University of Southampton(南安普顿大学电子与计算机科学学院)
  • School of Electrical Engineering and Computer Science, Queensland University of Technology(昆士兰科技大学电气工程与计算机科学学院)
  • School of Computer Science and Electronic Engineering, University of Essex(埃塞克斯大学计算机科学与电子工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑