arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoMix: 通过全局上下文和多检测器训练实现无描述符视觉定位

GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training

Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala

arXiv 2607.02486首次发表:更新:

发表机构

Aalto University; Tampere University; University of Oulu(阿尔托大学; 坦佩雷大学; 奥卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出GeoMix框架,通过局部方向与距离嵌入、全局上下文节点和多检测器混合训练,增强几何判别性,显著提升无描述符视觉定位精度。

AI 中文摘要

无描述符视觉定位消除了高维描述符存储,保护了场景隐私,并简化了地图维护,但其精度仍远落后于基于描述符的流程。我们发现这一差距源于纯几何匹配中几何判别性不足。没有视觉外观,当前方法未能充分利用局部几何线索,缺乏关键点间的全局上下文,并且过拟合于单一关键点检测器。我们进一步观察到,无描述符匹配自然支持多检测器训练,因为异构关键点可以在共享的纯几何空间中进行优化,而无需对齐描述符空间。基于这些见解,我们提出了GeoMix,一个无描述符的2D-3D匹配框架,在三个层面增强几何判别性。在局部层面,方向和距离感知嵌入通过细粒度空间结构丰富了邻域聚合。在全局层面,可学习的上下文节点通过交叉注意力聚合和重新分配场景范围的信息,以解决局部感受野之外的歧义。在训练层面,混合训练利用这种与检测器无关的几何空间,学习跨多个关键点检测器的表示。在MegaDepth、Cambridge Landmarks、7Scenes和Aachen Day-Night上的大量实验表明,GeoMix在无描述符方法中达到了新的最佳水平,将第75百分位的旋转误差降低了89%,平移误差降低了高达90%,同时零样本泛化到未见过的检测器,并缩小了与基于描述符流程的差距。代码可在$\href{this https URL}{\text{this links}}$获取。

英文摘要

Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet its accuracy still lags far behind descriptor-based pipelines. We identify this gap to insufficient geometric discriminability in geometry-only matching. Without visual appearance, current methods underutilize local geometry cues, lack the global context among keypoints, and overfit to a single keypoint detector. We further observe that descriptor-free matching naturally enables multi-detector training, as heterogeneous keypoints can be optimized in a shared geometry-only space without aligning descriptor spaces. Building on these insights, we propose GeoMix, a descriptor-free 2D-3D matching framework that strengthens geometric discriminability at three levels. Locally, directional and distance-aware embeddings enrich neighborhood aggregation with fine-grained spatial structure. Globally, learnable context nodes aggregate and redistribute scene-wide information via cross-attention to resolve ambiguities beyond local receptive fields. At the training level, Mix-Training exploits this detector-agnostic geometry space to learn representations across multiple keypoint detectors. Extensive experiments on MegaDepth, Cambridge Landmarks, 7Scenes, and Aachen Day-Night show that GeoMix sets a new state of the art among descriptor-free methods, reducing 75th-percentile rotation error by 89\% and translation error by up to 90\% over the previous best, while generalizing zero-shot to unseen detectors and narrowing the gap to descriptor-based pipelines. Code is available at $\href{https://github.com/YejunZhang/Geomix}{\text{this links}}$.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑