arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准检索几何:面向视觉地点识别的可靠性引导免训练聚合

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang

arXiv 2609.25937首次发表:更新:

发表机构

University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉地点识别中固定聚合抑制新环境区分性的问题,提出可靠性引导的免训练聚合方法TFA,利用跨码本一致性与谱统计校准检索,在多个基准上显著提升Recall@1。

AI 中文摘要

冻结的视觉基础模型为视觉地点识别提供可迁移特征,但固定的聚合方式可能抑制新环境中有用的区分信息。我们提出TFA,一种可靠性引导的免训练聚合方法,既不需要地点标签,也不需要特定任务权重更新。我们的关键观察是,可复现的检索不必具有区分性:独立的码本可以一致地检索少数数据库枢纽。TFA结合跨码本一致性、检索覆盖率和谱统计量,以控制残差分配、谱整形和全局特征融合。其谱核在零干预时精确恢复原始描述符相似性。仅数据库的TFA在访问查询前固定其规则;TFA-C64使用64张不重叠的无标签目标图像来校准后续查询的检索。在20个地面协议中,使用固定的DINOv2-B骨干和匹配分辨率,仅数据库的TFA在MSLS-val上将Recall@1比AnyLoc提升17.39个百分点,在SPED上提升9.55个百分点。C64缓解了驾驶环境中仅数据库校准的失败。在8个空中/跨视图协议中,TFA在16个DINOv2/DINOv3骨干-协议组合中的14个上取得最高的Recall@1,与所比较的免训练头部相比。在单独的原生系统比较中,基于DINOv2-G的TFA-C64在Pitts30k上达到91.46%的Recall@1,在VPAIR上达到76.29%,在所有五个基准上优于所展示的免训练比较器。这些结果表明,可靠性引导的聚合可以从冻结表示中恢复额外的检索能力,为地点监督稀缺的新环境提供实用基线。

英文摘要

Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-guided, training-free aggregation method requiring neither place labels nor task-specific weight updates. Our key observation is that reproducible retrieval need not be discriminative: independent codebooks can consistently retrieve a few database hubs. TFA combines cross-codebook agreement, retrieval coverage, and spectral statistics to control residual assignment, spectral shaping, and global-feature fusion. Its spectral kernel exactly recovers original descriptor similarity at zero intervention. Database-only TFA fixes its rules before accessing queries; TFA-C64 uses 64 disjoint unlabeled target images to calibrate retrieval for subsequent queries. Across 20 ground protocols with a fixed DINOv2-B backbone and matched resolution, database-only TFA improves Recall@1 over AnyLoc by 17.39 percentage points on MSLS-val and 9.55 on SPED. C64 mitigates failures of database-only calibration in driving environments. Across eight aerial/cross-view protocols, TFA achieves the highest Recall@1 among compared training-free heads in 14 of 16 DINOv2/DINOv3 backbone-protocol combinations. In a separate native-system comparison, DINOv2-G-based TFA-C64 reaches 91.46% Recall@1 on Pitts30k and 76.29% on VPAIR, outperforming the displayed training-free comparators on all five benchmarks. These results show that reliability-guided aggregation can recover additional retrieval capability from frozen representations, providing a practical baseline for new environments with scarce place supervision.

Comments26 pages, 5 figures, 9 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑