arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Glob3R:基于3D基础模型的全局运动恢复结构

Glob3R: Global Structure-from-Motion with 3D Foundation Models

Junyuan Deng, Heng Li, Kejie Qiu, Lingteng Qiu, Rui Peng, Weichao Shen, Weihao Yuan, Siyu Zhu, Zilong Dong, Ping Tan

arXiv 2607.09225首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Tongyi Lab, Alibaba Group; Nanjing University; Fudan University(香港科技大学; 阿里巴巴集团通义实验室; 南京大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对3D基础模型重建结果不准确及处理长序列或大图像集有缺陷的问题,提出Glob3R方法,通过增强主干、转换扭曲为特征轨迹、引入关联策略及进行全局优化等,实现强大准确重建,提升神经渲染质量。

AI 中文摘要

近期的3D几何基础模型,如VGGT,通过直接从输入图像预测相机姿态和3D场景点来提供强大的前馈3D重建。然而,其结果仍不准确,扩展到长序列或大型无序图像集通常需要逐块处理,这会引入漂移和不一致性。我们提出了Glob3R,一种基于3D基础模型的全局SfM风格重建方法。关键思想是显式优化前馈几何预测。为此,我们用一个轻量级的密集匹配头增强冻结的Pi3X主干,该头预测选定参考帧与相邻视图之间的图像扭曲。这些密集扭曲被转换为稀疏但可靠的多视图特征轨迹,为全局优化提供对应约束。我们进一步引入基于关键帧的滑动窗口关联策略,在重叠窗口中传播轨迹和相对姿态,实现可扩展的重建。最后,我们进行全局运动平均和束调整,以优化相机姿态,减少尺度不一致性,并恢复密集场景几何。在室内、室外、大规模驾驶和无序SfM基准上的大量实验表明,Glob3R实现了强大而准确的重建。它始终优于前馈基础模型基线和近期的可扩展重建方法,同时比经典SfM管道更稳健。优化后的姿态还带来了更高质量的神经渲染,验证了将基础模型先验与全局几何优化相结合的好处。

英文摘要

Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization. Project page: https://junyuandeng.github.io/Glob3r

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑