TAPVid-MV:跨多视图三维任意点跟踪基准
TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views
浏览论文内容
中文总结 AI 辅助
TAPVid-MV是首个针对相机运动下跨多视图长期三维点跟踪的基准,含多类场景数据,揭示几何恢复是三维点跟踪瓶颈,还支持多项相关任务。
中文摘要 AI 辅助
多相机系统在机器人技术、增强现实/虚拟现实(AR/VR)和自动驾驶领域的实用性日益提升,因为互补视图可减少深度歧义并在遮挡情况下保持可见性。然而,现有点跟踪基准聚焦于单一视频或静态多相机装置,均未测试相机运动下跨多个同步视图的长期三维点跟踪。我们推出TAPVid-MV(视频跨多视图任意点跟踪),首个针对该场景的基准。它包含精心筛选的284个序列、1142个校准相机流及109769个点轨迹,涵盖七个子集,涉及室内外场景,从机器人技术、人类活动到驾驶及合成程序场景。我们利用数据集特定的辅助模态获取这些轨迹:传感器深度、激光雷达(LiDAR)、同时定位与建图(SLAM)与运动恢复结构(SfM)点、人体网格、带姿态的物体网格及模拟数据。每个序列和轨迹均经人工标注者视觉验证。在超过30个基线方法中,无任何方法接近解决该任务。令人惊讶的是,现有多视图点跟踪器并未始终优于单目点跟踪器。通过在相同数据集上评估重建与点跟踪,TAPVid-MV有助于区分恢复几何结构的误差与点对应关系的误差。通过此联合分析,我们确定几何结构恢复是准确三维点跟踪的主要瓶颈。除多视图三维点跟踪外,我们发布的标注还支持单目二维与三维点跟踪、未来轨迹预测及四维(4D)重建。
英文摘要
Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusion. Existing point-tracking benchmarks, however, focus on a single video or static multi-camera rigs. None test long-term 3D point tracking across several synchronized views under camera motion. We introduce TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for this setting. It contains a curated set of 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks across seven subsets spanning indoor and outdoor domains, from robotics and human activity to driving and synthetic procedural scenes. We obtain these trajectories using dataset-specific auxiliary modalities: sensor depth, LiDAR, SLAM and SfM points, human meshes, posed object meshes, and simulation. Every sequence and trajectory is visually verified by human annotators. Across more than 30 baselines, no method comes close to solving the task. Surprisingly, existing multi-view point trackers do not consistently outperform monocular point trackers. By evaluating reconstruction and point tracking on the same datasets, TAPVid-MV helps distinguish errors in recovered geometry from errors in point correspondence. Through this joint analysis, we identify geometry recovery as a major bottleneck for accurate 3D point tracking. Beyond multi-view 3D point tracking, our released annotations support monocular 2D and 3D point tracking, future-trajectory prediction, and 4D reconstruction.
发表机构
- Google DeepMind(谷歌DeepMind)
- University College London(伦敦大学学院)
- University of Oxford(牛津大学)
- ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。