超越图像平面:用于多目标跟踪的世界 grounding 查询
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
浏览论文内容
中文总结 AI 辅助
针对单目视频多目标跟踪受图像平面模糊深度与空间关系的局限,提出端到端多目标跟踪器 PLANET,通过将二维跟踪数据集提升至三维并构建世界 grounding 查询等方法,在三个基准上取得最优性能。
中文摘要 AI 辅助
单目视频将三维场景记录为二维图像平面投影序列,会模糊深度与空间关系。多目标跟踪器主要仅利用图像平面中观测到的外观与几何信息来定位和关联目标,继承了这些模糊性。为解决这一局限,我们推出 PLANET——一种端到端多目标跟踪器,旨在超越图像平面。作为使能步骤,我们将现有二维跟踪数据集提升至三维。随后,我们通过将重建的三维场景几何嵌入查询形成过程中使用的特征与位置编码,构建世界 grounding 查询。辅助的三维位置预测任务进一步促使查询在训练期间编码目标位置。互补的双分辨率时间记忆则在更长时间间隔内保留该证据。最终,PLANET 在三个不同基准上实现了最优性能。
英文摘要
Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance and geometry observed only in the image plane, inheriting these ambiguities. To address this limitation, we introduce PLANET, an end-to-end multi-object tracker designed to move beyond the image plane. As an enabling step, we lift existing 2D tracking datasets into 3D. We then form world-grounded queries by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation. An auxiliary 3D location prediction task further encourages the queries to encode object positions during training. A complementary dual-resolution temporal memory preserves this evidence across longer temporal gaps. As a result, PLANET achieves state-of-the-art performance across three diverse benchmarks.
发表机构
- NVIDIA(英伟达)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。