发表机构
Google DeepMind; Simon Fraser University; University of Toronto(谷歌DeepMind; 西蒙菲莎大学; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有SfM方法在复杂视频中失效且缺乏对应可靠基准的问题,提出基于360°视频的ORBIT基准,实验显示多种SfM方法在该基准上表现不佳,为相关研究提供了测试平台。
AI 中文摘要
运动恢复结构(Structure-from-Motion,SfM)是三维感知的基石,但当前方法在应用于包含挑战性相机运动或动态场景的复杂视频时常常失效。更严重的是,该领域缺乏针对这类困难场景的可靠真值基准,难以衡量实际进展或确定最需要改进的方向。为解决这一缺口,我们提出了一个用于评估相机位姿估计的新基准。我们的核心见解是利用在线全景360°视频作为数据源,构建具有挑战性的片段,同时仍能实现稳健的真值轨迹恢复。这些视频的全景特性为跟踪相机运动提供了更丰富的视觉上下文,即使在部分视图受模糊、运动或动态物体影响时也能发挥作用。在跟踪完整360°视频中的相机运动后,我们裁剪并重新投影选定的部分,生成作为基准的透视视图片段,将其命名为ORBIT。实验表明,COLMAP以及近期基于优化和前馈的SfM方法在我们的基准上难以准确估计相机位姿。因此,ORBIT为研究人员提供了一个宝贵的测试平台,可在此平台上有意义地衡量在真正具有挑战性的实际SfM问题上的进展。
英文摘要
Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes. Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress or to pinpoint where improvements are most needed. To address this gap, we introduce a new benchmark for evaluating camera pose estimation. Our key insight is to leverage online panoramic 360° video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery. The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects. After tracking camera motion across full 360° videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark, called ORBIT. Experiments show that COLMAP, as well as recent optimization-based and feed-forward SfM methods struggle to accurately estimate camera poses on our benchmark. Hence, ORBIT provides a valuable testbed where researchers can meaningfully measure progress on truly challenging, real-world SfM problems.
CommentsA revision was Accepted at CVPR 2026