超越综述:视觉多目标跟踪中检测与关联的系统性实证研究
Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT
- GIST(光州科学技术院)
- University of Washington Tacoma(华盛顿大学塔科马分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过系统实验量化检测与关联组件对视觉多目标跟踪性能的贡献,发现检测质量主导性能,优于关联策略,并为系统设计提供实用指导。
AI中文摘要:
本文对最先进的多目标跟踪算法进行了全面的实验评估和详细分析,重点在于量化检测和关联组件对整体跟踪性能的各自贡献。与现有主要提供跟踪方法理论分类或类别划分的综述不同,我们的工作采用基于公开实现的严谨实验视角,为研究人员和从业者在方法选择和系统设计方面提供实用指导。我们引入了一个统一的流程示意图,整合了视觉多目标跟踪两大主要分支(基于检测的跟踪和端到端深度学习范式)的核心组件,并系统分析了目标检测、特征提取和数据关联模块。通过在标准基准(包括MOT16、MOT17、MOT20、SportsMOT、DanceTrack和CrowdTrack数据集)上的广泛实证研究,我们揭示了关键见解:(1)检测质量主导关联策略的性能,检测器的改进带来超过10%的提升,而精细化的关联策略提升不足5%;(2)现代深度学习检测器与专用重识别模型相结合,显著优于联合检测和嵌入方法;(3)基于Transformer的端到端方法对检测质量变化表现出更强的鲁棒性,但计算成本显著更高。我们通过大量实验得出的发现,提供了对MOT中组件级效应的关键见解,特别是检测质量相对于关联的主导影响,同时为在不同性能和鲁棒性要求下设计和优化MOT系统提供了实用见解。代码和实验设置可在该http URL获取。
英文摘要:
This paper presents a comprehensive experimental evaluation and detailed analysis of state-of-the-art multi-object tracking algorithms, with an emphasis on quantifying the individual contributions of detection and association components to overall tracking performance. Unlike existing surveys that primarily offer theoretical categorizations or taxonomies of tracking methods, our work adopts a rigorous experimental perspective grounded in publicly available implementations, providing practical guidance for researchers and practitioners in method selection and system design. We introduce a unified pipeline diagram that consolidates the core components across the two main branches of visual multi-object tracking: tracking-by-detection and end-to-end deep learning paradigms, and systematically analyze the object detection, feature extraction, and data association modules. Through extensive empirical studies on standard benchmarks, including MOT16, MOT17, MOT20, SportsMOT, DanceTrack, and CrowdTrack datasets, we reveal critical insights: (1) detection quality dominates association strategy performance, with detector improvements yielding more than 10% gains compared to less than 5% from refined association strategies; (2) modern deep learning detectors paired with specialized re-identification models significantly outperform joint detection and embedding approaches; and (3) transformer-based end-to-end methods exhibit greater robustness to detection quality variations but at a substantial computational cost. Our findings from extensive experiments provide key insights into component-level effects in MOT, particularly the dominant influence of detection quality relative to association, while offering practical insights for designing and optimizing MOT systems under varying performance and robustness requirements. Code and experimental setups are available at github.com/linh-gist/VisualMOT.