AI 中文总结
研究多目标跟踪中如何通过解决激活稀疏偏好问题,利用脉冲神经网络实现高效跟踪。核心方法是提出SpikingMOT,将轨迹状态分解并用预测误差校准后验。主要贡献是取得领先性能,减少参数和能量,为高效跟踪开辟新方向。
AI 中文摘要
多目标跟踪(MOT)在视觉感知中起着基础性作用,准确的轨迹预测对于复杂运动模式下可靠的目标关联至关重要。近期的跟踪器通过密集激活的人工神经网络改进了运动建模,但很大程度上忽略了这种密集响应对于轨迹预测是否必要。本文通过解决两个关键问题来制定激活稀疏偏好(ASP):一是如何识别能恰当且正式解释ASP的模型架构,二是如何将这种解释转化为有竞争力的跟踪性能。理论分析表明,在相同激活率下,稀疏门控并不比与状态无关的随机失活差。基于此,提出了SpikingMOT,一种基于脉冲驱动的跟踪器,它用脉冲神经网络(SNNs)自适应地对稀疏轨迹动态进行建模。具体而言,SpikingMOT将每个轨迹状态分解为伪轨迹基,并使用当前预测误差来校准下一帧预测的后验。通过这种受大脑启发的循环,SpikingMOT在大量实验中取得了领先性能,在SportsMOT上达到74.9 HOTA,在DanceTrack上达到56.5 HOTA,同时分别将参数和能量减少了72%和86.7%。这些结果将SNNs引入MOT,为高效跟踪开辟了一个有前景的方向。
英文摘要
Multi-object tracking (MOT) plays a fundamental role in visual perception, where accurate trajectory prediction is essential for reliable target association under complex motion patterns. Recent trackers have improved motion modeling with densely activated artificial neural networks, yet they largely overlook whether such dense responses are necessary for trajectory prediction. In this paper, we formulate activation sparsity preference (ASP) by tackling two key questions: 1. How can we identify a model architecture that appropriately and formally explains ASP, and 2. How can we translate this explanation into competitive tracking performance. Theoretical analysis shows that sparse gating is no worse than state-independent dropout under the same activation rate. Based on this insight, SpikingMOT is proposed as a spike-driven tracker that adaptively models sparse trajectory dynamics with spiking neural networks (SNNs). Specifically, SpikingMOT decomposes each trajectory state into pseudo-trajectory bases and uses the current prediction error to calibrate the posterior for next-frame prediction. With this brain-inspired loop, SpikingMOT achieves state-of-the-art performance in extensive experiments, 74.9 HOTA on SportsMOT and 56.5 HOTA on DanceTrack, while reducing the parameters and energy by 72% and 86.7%, respectively. These results bring SNNs into MOT, opening a promising direction for efficient tracking.