基于无检测器特征匹配与多视图轨迹优化的空地图像鲁棒运动恢复结构
Robust structure from motion for aerial-ground images via detector-free feature matching and multi-view track refinement
浏览论文内容
中文总结 AI 辅助
本研究提出结合无检测器匹配网络与多视图轨迹优化的方法,用于解决空地图像三维重建中特征匹配鲁棒性问题,在5°位姿误差下AUC较LoFTR提升93.9%,实现更高精度的ISfM重建。
中文摘要 AI 辅助
从空地图像进行集成式三维重建是生成高精度城市三维模型的关键,但视点、尺度和旋转的剧烈变化使得鲁棒特征匹配极具挑战性。为解决这些局限,本研究提出一种旋转鲁棒的无检测器匹配网络,结合多视图轨迹优化用于增量式运动恢复结构(ISfM)。所提工作流程包含四个关键模块:第一,旋转感知特征提取,用全向状态空间块(OSS Block)替代传统卷积,该块可沿八个对称方向选择性扫描,以建模长程空间依赖并合成旋转不变特征图;第二,多尺度注意力变换,利用四叉树注意力构建分层令牌金字塔,分离高关联令牌区域并丢弃无关区域,以线性计算复杂度捕获长程上下文;第三,双向特征匹配,执行对称的粗到精匹配方案,其中粗对齐在互最近邻约束下计算双向Softmax置信矩阵,精对齐则使用多层感知机回归亚像素坐标偏移;第四,多视图轨迹优化,采用集成索引结构评估局部空间邻近度,并将不相交的子轨迹链接到置信度最高的锚点,确保ISfM流程中特征的稳定可重复性。通过真实空地数据集,实验结果表明,与LoFTR相比,所提方法在5°位姿误差下的AUC提升了93.9%,且在ISfM重建中达到最高精度,精度提升幅度为27.6%至32.7%。该方法为空地图像的集成式三维重建提供了可靠解决方案。
英文摘要
Integrated 3D reconstruction from aerial-ground images is essential for generating high-precision urban 3D models, yet severe variations in viewpoint, scale, and rotation make robust feature matching highly challenging. To address these limitations, this study introduces a rotation-robust detector-free matching network coupled with multi-view track refinement for incremental Structure from Motion (ISfM). The proposed workflow features four key modules. First, rotation-aware feature extraction replaces traditional convolutions with an Omnidirectional State Space Block (OSS Block) that selectively scans across eight symmetrical directions to model long-range spatial dependencies and synthesize rotation-invariant feature maps. Second, multi-scale attention transformation utilizes quadtree attention to build a hierarchical token pyramid that isolates high-association token regions and discards irrelevant areas, capturing long-range context with linear computational complexity. Third, bi-directional feature matching executes a symmetric coarse-to-fine matching scheme where coarse alignment computes dual-direction Softmax confidence matrices under mutual nearest neighbor constraints, and fine alignment uses a multi-layer perceptron to regress sub-pixel coordinate offsets. Finally, multi-view track refinement employs an integrated indexing structure to evaluate localized spatial proximity and link disjoint sub-tracks to the highest-confidence anchor point, ensuring stable feature repeatability across the ISfM pipeline. By using real aerial-ground datasets, experimental results demonstrate that the proposed method improves AUC at 5° pose error by 93.9% compared with LoFTR and achieves the highest precision in ISfM reconstruction, with the improved accuracy ranging from 27.6% to 32.7%. The proposed method provides a reliable solution for integrated 3D reconstruction of aerial-ground images.
发表机构
- School of Architecture and Urban Planning, Shenzhen University(深圳大学建筑与城市规划学院)
- Guangdong Key Laboratory of Urban Informatics, Shenzhen University(深圳大学广东省城市信息学重点实验室)
- MNR Key Laboratory for Geo-Environmental Monitoring of Great Bay Area, Shenzhen University(深圳大学自然资源部大湾区地理环境监测重点实验室)
机构由 AI 辅助整理,请以论文原文为准。