arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SGFormer:用于鲁棒局部特征匹配的结构引导Transformer

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Zhihua Xu, Runyu Zhu, Rongjun Qin

arXiv 2608.03423首次发表:更新:

AI 中文总结

针对局部特征匹配中Transformer注意力发散的问题,提出SGFormer结构引导Transformer,通过三重结构注意力模块增强显著结构区域的注意力,有效提升了匹配精度。

AI 中文摘要

局部特征匹配是摄影测量学的基础组成部分,可为3D重建、立体测绘和视觉定位等任务提供关键的精确图像对应关系。尽管近期无检测器的匹配方法(如LoFTR)推动了该领域的发展,但利用无约束注意力机制的全局范围建模能力所获得的全局特征,会使模型在某些场景中对显著结构的注意力受到损害。这种局限性导致了我们定义为“注意力发散”的现象:一部分高置信度匹配分布在有效匹配区域(重叠区域)之外,尤其是在视角变化大的场景中。这是因为标准Transformer中,无关区域的相似特征可能获得相同的权重并被考虑,从而限制了摄影测量挑战性环境中的匹配可靠性。为解决特征匹配中的该问题,我们提出SGFormer(Structure-Guided Transformer,结构引导Transformer),这是一种新颖的结构感知匹配网络,可自适应更新重叠区域内显著结构附近特征的注意力。SGFormer采用半密集的由粗到细流程,并在主干网络中融入所提出的三重结构注意力(Triple-Structure-Attention,TSA)模块以提取有区分度的特征。TSA模块利用网络早期层的浅层局部特征来增强显著结构周围的表示,引导后续Transformer阶段在全局范围内强化模型对显著结构区域的关注。因此,SGFormer加强了对视觉一致区域的注意力,同时减轻了非重叠区域的影响。大量实验表明,SGFormer显著缓解了注意力发散问题,并提高了匹配精度。

英文摘要

Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios. This limitation leads to a phenomenon we define as attention divergence, wherein a portion of high-confidence matches are distributed outside the valid matching region (overlapping region), especially in scenes with large viewpoint variations. This occurs because similar features in irrelevant regions may receive equal weighting and consideration within the standard Transformer, limiting matching reliability in challenging photogrammetric environments. To address this issue in feature matching, we propose SGFormer (Structure-Guided Transformer), a novel structure-aware matching network that adaptively updates attention on features near salient structure in overlapping regions. SGFormer employs a semi-dense coarse-to-fine pipeline and incorporates the proposed Triple-Structure-Attention (TSA) module into the backbone net for extracting distinctive features. The TSA module utilizes shallow local features from early network layers to enhance the representation around salient structure, guiding subsequent transformer stages to intensify the model's focus on regions with salient structure across the global scope. SGFormer, thereby reinforcing attention to visually consistent areas while mitigating the influence of non-overlapping regions. Extensive experiments show that SGFormer significantly mitigates attention divergence and improves matching accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑