发表机构
University of Belgrade(贝尔格莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统地将超参数优化应用于多目标跟踪,提出多保真度贪心坐标搜索(MFGCS),在八种组合中优于手动调参和已发表结果,最高提升16.05 HOTA点。
AI 中文摘要
多目标跟踪(MOT)以基于检测的跟踪范式为主导,该类方法通常依赖少量超参数,而这些超参数传统上由人工选择。调整这些超参数需要反复进行专家指导的实验,而用于选择报告值的流程往往未经系统评估或完整记录。超参数优化(HPO)使这一过程自动化,但在MOT中仍很少使用,且现有的将HPO应用于MOT的研究早于基于现代深度检测器的跟踪器和HOTA评估。我们系统地将HPO应用于两个数据集和四种基于检测的跟踪方法。我们还提出了多保真度贪心坐标搜索(MFGCS),该方法每次优化一个超参数,首先在小部分场景子集上评估候选值,仅在完整数据集上重新评估有希望的候选值。在所有八种跟踪器-数据集组合中,树结构Parzen估计器(TPE)和MFGCS均优于我们手动调整的配置和相应的已发表结果,分别最多提升4.38和16.05个HOTA点。MFGCS在八种组合中的七种中比TPE更快达到预定的HOTA目标。在每个跟踪器-数据集对中,所有优化器共享相同的搜索空间和评估流程,从而隔离了搜索策略的影响。我们发布代码和调整后的配置,以使未来的工作能够与系统优化而非默认或手动调整的基线进行比较。
英文摘要
Multi-object tracking (MOT) is dominated by the tracking-by-detection paradigm, whose methods typically rely on a small set of hyperparameters that are conventionally chosen by hand. Tuning them requires repeated expert-guided experimentation, while the procedures used to select reported values are often not systematically evaluated or fully documented. Hyperparameter optimization (HPO) automates this process, yet it remains rarely used in MOT, and existing studies applying HPO to MOT predate modern deep-detector-based trackers and HOTA evaluation. We systematically apply HPO across two datasets and four tracking-by-detection methods. We also propose Multi-Fidelity Greedy Coordinate Search (MFGCS), which optimizes one hyperparameter at a time by first evaluating candidate values on a small subset of scenes and re-evaluating only promising candidates on the full dataset. Across all eight tracker-dataset combinations, the Tree-structured Parzen Estimator (TPE) and MFGCS outperform both our hand-tuned configurations and the corresponding published results, with improvements of up to 4.38 and 16.05 HOTA points, respectively. MFGCS also reaches a predefined HOTA target faster than TPE in seven of the eight combinations. Within each tracker-dataset pair, all optimizers share the same search space and evaluation pipeline, isolating the effect of the search strategy. We release the code and tuned configurations to enable future work to compare against systematically optimized rather than default or manually tuned baselines.
Comments23 pages, 7 figures, 12 tables