发表机构
RWTH Aachen University(亚琛工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ARROW是一种前馈模型,通过顺序不变查询方法统一任意图像集的3D重建与点跟踪,在WorldTrack和TAPVid-3D上达到最先进水平,并优于专用多视图跟踪器。
AI 中文摘要
动态场景可能由移动相机、多视频流或在不同时间拍摄的图像捕捉。这些观测揭示了场景几何和运动的互补方面,但将它们整合需要建立跨视点、拍摄时间和可见性变化的对应关系。我们引入ARROW,一种前馈模型,统一了来自任意图像集的3D重建和3D点跟踪。其核心是一种新颖的顺序不变查询方法,允许查询与任意输入中的观测相关联。我们表明,在训练期间向模型暴露更多样化的输入集可提高任务性能。此外,所得模型能够泛化到更广泛的任务,包括多视图跟踪。采用此策略训练后,ARROW在WorldTrack和TAPVid-3D上的3D跟踪中建立了新的最先进水平,并在改编的仅RGB MVTracker基准上超越了专用的多视图跟踪器,同时在3D重建任务中保持竞争力。代码和权重公开可用。
英文摘要
Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks. Code and weights are publicly available.
CommentsProject page at: https://www.vision.rwth-aachen.de/arrow