发表机构
The City College of New York; INSAIT; Sofia University “St. Kliment Ohridski”; Julius-Maximilians-Universität Würzburg; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Shenzhen University of Advanced Technology(纽约城市学院; INSAIT; 圣克莱门特奥赫里德斯基索非亚大学; 维尔茨堡大学; 中国科学院计算技术研究所人工智能安全国家重点实验室; 中国科学院大学; 深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对统一多模态跟踪模型计算负担大的问题,提出双对齐蒸馏框架,结合知识蒸馏与结构剪枝,在五组基准上实现了高准确率与5倍速度提升的紧凑模型。
AI 中文摘要
统一多模态目标跟踪通过利用互补的传感器数据(如RGB、热成像、深度数据)已实现了出色的鲁棒性,但最先进模型沉重的计算负担阻碍了它们在资源受限的边缘设备上部署。在本研究中,我们发现预测头是一个关键却常被忽视的效率瓶颈。通过系统性精简解码器架构,我们释放了实时推理的潜力,但同时在轻量化学生模型与重型教师模型之间引入了容量差距。为解决这一问题,我们对17种蒸馏策略进行了系统分析,并提出了双对齐蒸馏框架。我们的核心见解是,有效的压缩需要将知识传递解耦为两个互补流:(1)空间表示对齐,采用特征蒸馏来强化学生模型对前景目标的空间聚焦(“跟踪位置”);(2)语义分布对齐,利用基于logit的蒸馏来对齐决策边界并传递判别性暗知识(“跟踪内容”)。在五个基准上的大量实验表明,我们的方法显著优于复杂的最先进方法。值得注意的是,我们的蒸馏模型在RGBT234数据集上达到了91.5%的MPR,在单块RTX 4090上以54 FPS运行,实现了比教师模型快5倍的速度,同时保持了更优的准确率。
英文摘要
Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy computational burden of state-of-the-art models hinders their deployment on resource-constrained edge devices. In this work, we identify the prediction head as a critical but often overlooked efficiency bottleneck. By strategically streamlining the decoder architecture, we unlock the potential for real-time inference but simultaneously introduce a capacity gap between the lightweight student and the heavy teacher. To resolve this, we conduct a systematic analysis of 17 distillation strategies and introduce a Dual-Alignment Distillation framework. Our key insight is that effective compression requires decoupling knowledge transfer into two complementary streams: (1) Spatial Representation Alignment, which employs feature distillation to sharpen the student's spatial focus on foreground targets ("Where to track"); and (2) Semantic Distribution Alignment, which utilizes logit-based distillation to align decision boundaries and transfer discriminative dark knowledge ("What to track"). Extensive experiments across five benchmarks demonstrate that our approach significantly outperforms complex state-of-the-art methods. Notably, our distilled model achieves 91.5% MPR on RGBT234 and operates at 54 FPS on a single RTX 4090, representing a 5x speedup over the teacher model while maintaining superior accuracy.