基于循环稀疏查询的ID预测在线多相机3D跟踪
Online Multi-Camera 3D Tracking via ID Prediction over Recurrent Sparse Queries
浏览论文内容
中文总结 AI 辅助
提出在线多相机3D跟踪架构,通过循环稀疏查询显式预测ID,结合Sparse4D检测与MOTIP解码,在AI City Challenge上HOTA从29.63提升至38.01,排名第三。
中文摘要 AI 辅助
在线多相机3D跟踪必须在同步视图中维护场景全局身份,然而基于查询的跟踪器仅将这些身份隐式地保存在实例库中,当查询中断时,这些身份会碎片化。我们提出了一种在线架构,通过显式地在循环稀疏查询上预测ID来恢复关联准确性。一个由外而内的Sparse4D检测器将校准视图融合为世界坐标系下的3D检测结果,同时传播一个稀疏查询库,并且一个因果MOTIP ID解码器将检测结果与有限轨迹记忆进行关联。我们调整了MOTIP的相对ID预测和回收槽位运行时,以适用于全局融合的3D观测,并引入了度量空间门控和基于邻近性的新生恢复机制。在官方2026年AI城市挑战赛赛道1测试集上,我们的方法将HOTA从使用原生实例库身份的29.63提升至38.01,主要得益于AssA从20.83提升至31.10,并在公开排行榜上排名第三。对每个场景全部9,000帧进行的全序列验证表明,解耦的ID训练相比原生身份能提升HOTA,而将检测器训练与分离的ID目标同时进行则会产生依赖于场景的收益和损失。
英文摘要
Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only implicitly in the instance bank, where they fragment upon query interruption. We present an online architecture that recovers association accuracy by predicting IDs explicitly over recurrent sparse queries. An outside-in Sparse4D detector fuses calibrated views into world frame 3D detections while propagating a sparse query bank, and a causal MOTIP ID decoder associates detections against a finite trajectory memory. We adapt MOTIP's relative-ID prediction and recycled slot runtime to globally fused 3D observations, and introduce metric spatial gating and proximity based newborn recovery. On the official 2026 AI City Challenge Track 1 test set, our method raises HOTA from 29.63 with native instance bank identities to 38.01, primarily through an AssA increase from 20.83 to 31.10, and ranks third on the public leaderboard. Full-sequence validation over all 9,000 frames of each scene shows that decoupled ID training improves HOTA over native identities, whereas continuing detector training alongside the detached ID objective produces scene-dependent gains and losses.
发表机构
- Playbox Inc.(Playbox 公司)
机构由 AI 辅助整理,请以论文原文为准。