发表机构
Glasgow College, University of Electronic Science and Technology of China; Key Laboratory for Urban Habitat Environmental Science and Technology, School of Environment and Energy, Peking University Shenzhen Graduate School; School of Management Science and Real Estate, Chongqing University(电子科技大学格拉斯哥学院; 北京大学深圳研究生院环境与能源学院城市人居环境科学与技术重点实验室; 重庆大学管理科学与房地产学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多目标跟踪中对象外观相似时身份保持难的问题,提出VLA-ReID方法,将重新识别建模为视频级关联,通过聚合历史轨迹特征等进行全局关联优化,实验显示该方法有效提升多项指标并减少身份切换。
AI 中文摘要
多目标跟踪旨在在视频中定位多个对象并随时间保持其身份。在诸如蜂群场景中,当对象小、密集分布且外观高度相似时,长期身份保持仍然困难。现有跟踪器依赖通过单实例分配训练的重新识别模型。然而,在推理时,多目标跟踪需要多个轨迹和检测之间的全局分配,这导致训练与推理不匹配,会造成视觉相似对象间的身份切换。现有方法还常需大量额外注释。我们提出视频级关联重新识别(VLA-ReID),将重新识别重新表述为视频级关联建模。它使用聚合的历史轨迹特征作为查询,当前帧的所有检测作为候选,在每一帧直接优化全局关联。此外,帧共同外观估计(FCAE)从当前帧检测中估计共同外观方向,共同外观抑制(CAS)从轨迹和检测特征中去除沿此方向的相应分量,无需额外注释放大高度相似对象间的差异。在BEE24上的实验表明,VLA-ReID比现有跟踪器在HOTA上提高1.1、MOTA提高0.3、AssR提高2.6、AssA提高0.7、IDF1提高0.8,同时身份切换减少28%。这些结果证明了视频级重新识别建模在基于外观的多目标跟踪关联中的有效性。
英文摘要
Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time. Long-term identity preservation remains difficult when objects are small, densely distributed, and highly similar in appearance, as in bee swarm scenes. Existing trackers rely on re-identification (re-ID) models trained through single-instance assignment (instance-level querying). At inference, however, MOT requires global assignment between multiple trajectories and detections, corresponding to video-level querying. This training-inference mismatch can cause identity switches among visually similar objects. Existing approaches also often require substantial additional annotations to enhance appearance discrimination. We propose Video-Level Association re-ID (VLA-ReID), which reformulates re-ID as video-level association modeling. It uses aggregated historical trajectory features as queries and all current-frame detections as candidates, enabling direct optimization of their global association at each frame. In addition, Frame-Common Appearance Estimation (FCAE) estimates a common appearance direction from current-frame detections, while Common-Appearance Suppression (CAS) removes the corresponding component along this direction from trajectory and detection features. This amplifies discriminative differences among highly similar objects without additional annotations. Experiments on BEE24 show that VLA-ReID improves HOTA by 1.1, MOTA by 0.3, AssR by 2.6, AssA by 0.7, and IDF1 by 0.8 over state-of-the-art trackers, while reducing identity switches by 28%. These results demonstrate the effectiveness of video-level re-ID modeling for appearance-based association in MOT.
Comments11 pages, 8 figures, 2 tables