MOTIP2:用于端到端多目标跟踪的空间先验
MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking
- École Centrale de Lyon(里昂中央理工学院)
- IDEMIA
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对端到端多目标跟踪中的空间不合理错误,提出三个空间先验(数据、损失、表示),构建MOTIP2模型,在多个基准上达到新最先进水平,并支持速度-精度权衡。
AI中文摘要:
端到端多目标跟踪器在关联困难的基准测试上缩小了与经典跟踪-检测方法的差距。然而,它们仍然会犯经典跟踪器不会犯的空间上不合理的错误,例如将同一身份分配给帧中相对两侧的物体。模型可以学习避免这些错误,但跟踪标注稀缺,因此我们改为显式编码空间先验,同时保持推理完全端到端,无需事后关联。我们提出了三个空间先验,分别作用于数据、损失和表示阶段。空间身份切换(Spatial ID Switches)将轨迹排列偏向于空间重叠的物体,减少训练与推理混淆之间的不匹配。空间身份损失(Spatial ID Loss)根据每个身份的边界框距离缩放其惩罚,因此远距离切换比近距离切换代价更高。空间锚点(Spatial Anchor)为每个轨迹令牌提供其帧位置,作为注意力的显式空间线索。我们将这三个先验实例化到MOTIP2中,这是一个从MOTIP改编并基于实时DEIM检测变换器的跟踪器。在没有额外数据训练的情况下,其主模型MOTIP2-L在DanceTrack上达到73.4 HOTA,在SportsMOT上达到76.0,在PersonPath22上达到71.1 IDF1,创下新的最先进水平。MOTIP2是一个覆盖速度-精度权衡的模型系列:一个更轻的模型MOTIP2-S在超过3倍速度下匹配原始MOTIP,而MOTIP2-X在DanceTrack上达到74.8 HOTA。
英文摘要:
End-to-end multi-object trackers have narrowed the gap with classical tracking-by-detection on association-difficult benchmarks. Yet they still make spatially implausible errors no classical tracker would, such as assigning one identity to objects on opposite sides of the frame. A model could learn to avoid them, but tracking annotations are scarce, so we encode spatial priors explicitly instead, while keeping inference fully end-to-end with no post-hoc association. We propose three spatial priors, at the data, loss, and representation stages. Spatial ID Switches bias trajectory permutations toward spatially overlapping objects, reducing the mismatch between training and inference confusions. Spatial ID Loss scales each identity's penalty by its box distance, so a distant switch costs more than a nearby one. Spatial Anchor gives each track token its frame position, an explicit spatial cue for attention. We instantiate the three priors in MOTIP2, a tracker adapted from MOTIP and built on the real-time DEIM detection transformer. Trained without extra data, its main model, MOTIP2-L, sets a new state of the art: 73.4 HOTA on DanceTrack, 76.0 on SportsMOT, and 71.1 IDF1 on PersonPath22. MOTIP2 is a family of models spanning the speed-accuracy trade-off: a lighter model, MOTIP2-S, matches the original MOTIP at over 3x the speed, and MOTIP2-X reaches 74.8 HOTA on DanceTrack.