发表机构
City University of Hong Kong; Dalian University of Technology(香港城市大学; 大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉跟踪中模板的权衡瓶颈,提出PATT师生训练框架,利用训练阶段的特权外观提升跟踪性能,移除教师组件后可部署,在七个基准上实现稳定增益。
AI 中文摘要
目标模板定义了视觉跟踪器的搜索对象,但推理时可用的模板在定位确定性与外观新鲜度之间存在权衡:初始真实模板准确但会过时,而近期模板更能反映当前外观,但来自不确定预测的裁剪。我们通过一个不可部署的“神谕”(oracle)量化了这一瓶颈,该神谕提供精确的当前帧目标裁剪,使LaSOT上的AUC提升15.2个百分点。这一差距揭示了仅训练阶段的机会:帧级真实值提供精确的当前和未来帧目标裁剪,尽管部署时无法获得此类裁剪。我们提出用于跟踪的特权外观迁移(Privileged Appearance Transfer for Tracking, PATT),这是一种师生训练框架,通过多级表示预测将这些特权外观迁移到可部署的跟踪器。特权教师观察过去、当前和未来帧的精确目标裁剪,而学生仅接收过去帧的模板,并学习预测教师的搜索表示。为避免迁移不可靠的教师信号,PATT通过教师相对于学生的相对定位优势及其绝对定位准确性对该迁移进行加权。训练完成后,移除教师、潜在预测器、可靠性权重和特权裁剪,仅保留标准的学生端推理。在两个模型规模的七个基准上,PATT在长期和短期跟踪协议下均实现了一致的性能提升。
英文摘要
Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better reflect the current appearance but are cropped from uncertain predictions. We quantify this bottleneck with a non-deployable oracle that supplies an exact current-frame target crop, improving AUC on LaSOT by 15.2 percentage points. This gap reveals a training-only opportunity: frame-level ground truths provide exact current- and future-frame target crops, although such crops are unavailable at deployment. We introduce Privileged Appearance Transfer for Tracking (PATT), a teacher-student training framework that transfers these privileged appearances to a deployable tracker through multi-level representation prediction. The privileged teacher observes exact target crops from past, current, and future frames, whereas the student receives only past-frame templates and learns to predict the teacher's search representations. To avoid transferring unreliable teacher signals, PATT weights this transfer by the teacher's relative localization advantage over the student and its absolute localization accuracy. After training, the teacher, latent predictor, reliability weights, and privileged crops are removed, leaving standard student-only inference. Across seven benchmarks at two model scales, PATT achieves consistent gains under both long- and short-term tracking protocols.
Comments13 pages, 2 figures