arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00925cs.CV

观察与回溯:用于全景SLAM的冻结基础模型中的隐式注意力与潜在方向

Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

Zhuang Xiong, Guohao Zhang, Chen Zhang, Zheyu Jiang, Yuchao Mei, Qingshan Xu, Wenbing Tao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对全景SLAM的相机倾斜、尺度漂移等问题,基于冻结全景几何基础模型的隐式线索提出HALO-SLAM,在五个基准的125个序列上实现100%序列成功率,ATE显著优于现有方法。

中文摘要 AI 辅助

单目全景SLAM得益于大相机旋转下的大量视觉重叠,但仍易受相机倾斜、尺度漂移和错误回环检测的影响。我们发现,冻结的全景几何基础模型除了显式几何输出外,还提供有用的内部线索:中间token编码相机坐标系中的重力,而跨视图注意力为潜在重访提供兼容性线索。基于这些线索,我们提出HALO-SLAM。重力读出实现了无IMU的球面正位标准化。对于回环检测,我们引入了一种成本感知的三级级联方法,结合DBoW2事件级检索、基于注意力的兼容性过滤,以及通过对称子图增强实现的密集几何验证。被接受的重访在两个局部度量中产生像素对齐的3D-3D对应关系,从中估计出鲁棒的Sim(3)约束,并与全局位姿图中的序列约束联合优化。在来自五个真实世界全景基准的125个序列上,我们的方法在规定标准下实现了100%的序列成功率(125/125),并且在所有五个基准上的ATE均为评估方法中最低,相比每个基准上最佳的ERP原生基线,ATE降低了30%-88%。

英文摘要

Monocular panoramic SLAM benefits from substantial visual overlap under large camera rotations, yet remains prone to errors caused by camera tilt, scale drift, and false loop closures. We show that a frozen panoramic geometry foundation model provides useful internal cues beyond its explicit geometric outputs: intermediate tokens encode gravity in the camera frame, while cross-view attention provides a compatibility cue for potential revisits. Building on these cues, we present HALO-SLAM. A gravity readout enables IMU-free spherical upright canonicalization. For loop closure, we introduce a cost-aware three-stage cascade combining DBoW2 event-level retrieval, attention-based compatibility filtering, and dense geometric validation through symmetric submap augmentation. Accepted revisits yield pixel-aligned 3D--3D correspondences in both local gauges, from which robust $\mathrm{Sim}(3)$ constraints are estimated and jointly optimized with sequential constraints in a global pose graph. Across 125 sequences from five real-world panoramic benchmarks, our method achieves \textbf{100\%} sequence success (\textbf{125/125}) under the stated criterion and the lowest ATE among the evaluated methods on all five benchmarks, reducing ATE by \textbf{30--88\%} relative to the best ERP-native baseline on each benchmark.

↑