arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04369cs.CV

AdaptVPR:用于鲁棒视觉地点识别的路线感知难正样本生成

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

Shunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu

首次发表
浏览论文内容

中文总结 AI 辅助

提出AdaptVPR框架生成难正样本,构建AdaptCities数据集,提升VPR在域偏移场景下的鲁棒性,R@1最高提升9.2%

中文摘要 AI 辅助

视觉地点识别(VPR)通过检索相同或邻近地点的数据库图像来定位查询图像,但其鲁棒性常因光照、天气、季节变化及动态遮挡引发的域偏移而下降。现有训练数据中同一地点的外观多样性有限是诱因之一。为解决该问题,我们提出AdaptVPR——一种路线感知的生成式数据增强框架,用于构建同一地点的难正样本以实现鲁棒VPR训练。AdaptVPR首先使用视觉语言模型解析场景属性并评估编辑可行性,同时基于可编辑性评分和风险约束,由基于规则的调度器确定生成路线。生成过程分解为三条互补路线:全局外观路线引入天气、光照、一天时段的全局场景变化;局部遮挡路线插入合理的动态遮挡物;双路线结合两类扰动以产生更具挑战性的外观偏移。每个生成的候选样本均通过面向VPR的验证方案评估,该方案基于几何一致性和外观多样性,可降低结构漂移风险同时确保足够的外观变化。全局候选样本生成一次,验证不通过则被丢弃;局部遮挡和双路线候选样本则利用验证反馈进行有限的提示优化与重新生成。借助该框架,我们构建了包含16万个经验证的合成同一地点难正样本的AdaptCities数据集。在多个VPR基线模型和视觉基础骨干网络上开展的实验显示,该框架在标准基准测试中取得了一致提升,在挑战性域偏移场景下实现了显著改进,R@1指标提升最高达9.2%。源代码和数据资源可在此URL公开获取。

英文摘要

Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby place, yet its robustness is often degraded by domain shifts arising from illumination, weather, seasonal changes, and dynamic occlusions. One contributing factor is the limited appearance diversity of the same place in existing training data. To address this issue, we propose AdaptVPR, a route-aware generative augmentation framework that constructs same-place hard positives for robust VPR training. AdaptVPR first uses a vision language model to parse scene attributes and estimate editing feasibility, while a rule-based scheduler determines the generation route according to editability scores and risk constraints. The generation process is decomposed into three complementary routes: the Global Appearance Route introduces global scene changes in weather, illumination, and time of day; the Local Occlusion Route inserts plausible dynamic occluders; and the Dual Route combines both types of perturbations to produce more challenging appearance shifts. Each generated candidate is evaluated using a VPR-oriented verification scheme based on geometric consistency and appearance diversity, reducing the risk of structural drift while ensuring sufficient appearance variation. Global candidates are generated once and rejected if verification fails, while Local Occlusion and Dual candidates use verification feedback for limited prompt refinement and regeneration. Using this framework, we construct AdaptCities, containing 160K verified synthetic same-place hard positives. Experiments across multiple VPR baselines and vision foundation backbones show consistent gains on standard benchmarks and substantial improvements under challenging domain shifts, with R@1 gains of up to 9.2%. The source code and data resources are publicly available at https://github.com/chenshunpeng/AdaptVPR.

发表机构

  • School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
  • Shandong Computer Science Center, Qilu University of Technology(齐鲁工业大学山东省计算中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑