TrackEverything:通过去重3D场景表示实现长时程密集追踪
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
- Carnegie Mellon University(卡内基梅隆大学)
- Meta
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TrackEverything通过将视频表示为世界坐标中的持久3D场景轨迹,结合体素化去重、端点与轨迹精化器及3D WAFT特征采样,在40GB内存内实现超1000帧全点密集追踪,短片段APD优于开源密集追踪器20%以上。
AI中文摘要:
现有的点追踪模型面临一个根本性的权衡:它们要么在长时程内追踪稀疏的查询点集,要么仅在短片段中追踪所有点。我们提出TrackEverything,一种3D点追踪器,通过将视频表示为世界坐标中的持久3D场景轨迹来打破这一权衡。基于视频是底层3D世界的2D投影这一洞见,TrackEverything将模型复杂度与视频时长解耦,使其能够随独特的物理场景几何结构扩展。我们的方法引入了三项关键创新。首先,我们在滑动窗口边界采用基于体素化的去重机制来合并共位轨迹,防止对同一表面的重复观测产生冗余累积。其次,我们将追踪分解为一个端点精化器(预测每个点的目的地及静态与动态分类)和一个轻量级轨迹精化器(仅对动态点解码密集轨迹)。第三,我们提出3D WAFT,用场景云中的高效特征采样替代内存受限的4D相关体积。据我们所知,TrackEverything是首个能够在40 GB GPU内存内追踪超过1000帧视频中所有可见点的3D追踪器。在TAPVid-3D上,TrackEverything在短片段上的APD(平均点距离)比所有开源的全帧密集3D追踪器高出20%以上,同时在长序列上与最先进的稀疏追踪器保持竞争力,尽管其追踪的点数远多于后者。
英文摘要:
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.