arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对查询传播多目标跟踪的片段级主动学习,利用轨迹状态扰动探测关联不稳定性

Probing Association Instability with Track-State Perturbations for Clip-Level Active Learning in Query-Propagation Multi-Object Tracking

Riku Inoue, Shogo Sato, Kazuhiko Murasaki, Tomoyasu Shimada, Toshihiko Nishimura, Ryuichi Tanida

arXiv 2608.17224首次发表:更新:

AI 中文总结

针对查询传播多目标跟踪的高标注成本问题,提出QPID方法,通过轨迹状态扰动探测关联不稳定性并结合视觉覆盖选择标注批次,在相同预算下优于主动学习基线。

AI 中文摘要

训练查询传播端到端多目标跟踪(MOT)模型需要跨视频序列的密集边界框和身份标注,这使得数据集构建成本高昂。片段级主动学习通过选择视频片段进行标注来降低该成本,但此前基于输出级时间不确定性的获取准则可能遗漏那些其信息性源于传播轨迹状态中关联不稳定性的片段。我们提出QPID(Query-Propagation Instability and Diversity,查询传播不稳定性与多样性),一种针对查询传播MOT的片段获取方法,其目标是传播轨迹状态中的关联不稳定性。QPID通过对内部轨迹状态施加双侧扰动并测量与干净参考分支的预测差异来估计这种不稳定性。核心思路是:在稳定片段中,每条传播轨迹在微小扰动下应继续跟踪同一目标;而在模糊片段中,轨迹状态的微小变化可改变该轨迹跟踪的目标,进而导致定位或置信度变化。QPID通过两个指标测量这些扰动诱导的预测差异:定位漂移(Localization Drift)和熵加权置信度差异(Entropy-Weighted Confidence Discrepancy)。这些指标被聚合成片段级关联不稳定性分数。为避免仅基于不确定性的冗余选择,QPID利用带轨迹级视觉原型的不确定性加权视觉覆盖,从高不稳定性片段中选择具有代表性的标注批次。在DanceTrack和SportsMOT数据集上,使用MeMOTR和SambaMOTR模型进行的实验表明,在相同标注预算下,QPID相较于主动学习基线方法实现了优异性能。

英文摘要

Training query-propagation end-to-end multi-object tracking (MOT) models requires dense bounding-box and identity annotations across video sequences, making dataset construction expensive. Clip-level active learning reduces this cost by selecting video clips for annotation, but prior acquisition criteria based on output-level temporal uncertainty may miss clips whose informativeness comes from association instability in propagated track states. We propose QPID (Query-Propagation Instability and Diversity), a clip acquisition method for query-propagation MOT that targets association instability in propagated track states. QPID estimates this instability by applying two-sided perturbations to internal track states and measuring prediction differences from a clean reference branch. The key idea is that, in stable clips, each propagated track should continue to follow the same target under small perturbations, whereas in ambiguous clips, small changes in the track state can alter which target the track follows, leading to changes in localization or confidence. QPID measures these perturbation-induced prediction differences with two metrics: Localization Drift and Entropy-Weighted Confidence Discrepancy. These metrics are aggregated into a clip-level association-instability score. To avoid redundant uncertainty-only selection, QPID selects a representative annotation batch from high-instability clips using Uncertainty-Weighted Visual Coverage with track-level visual prototypes. Experiments on DanceTrack and SportsMOT with MeMOTR and SambaMOTR show that QPID achieves strong performance compared with active learning baselines under the same annotation budget.

CommentsAccepted at the 37th British Machine Vision Conference (BMVC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑