arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FeedbackTrack:受视觉皮层启发的Transformer跟踪器跨帧反馈机制

FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking

Yueyang Cang, Xiaoteng Zhang, Zhiyuan Ning, Yuchen He, Li Shi

arXiv 2608.09369首次发表:更新:

AI 中文总结

该研究提出受视觉皮层启发的FeedbackTrack框架,为Transformer跟踪器引入跨帧反馈机制,在LaSOT、GOT-10k等数据集上提升跟踪性能,且参数增量极小。

AI 中文摘要

视觉目标跟踪需要有效的时间整合,但大多数Transformer跟踪器仍主要依赖前馈特征提取。现有时间机制通常更新模板、提示、查询或预测状态,而中间表示很少被重用以调制相应处理阶段。我们提出FeedbackTrack,一个受视觉皮层启发的框架,将稀疏的组级层对齐跨帧反馈引入预训练Transformer跟踪器。前一帧的中间状态被分离、缓存,并通过两条轻量路径返回至当前帧对应的Transformer组:用于令牌级查询调制的查询反馈,以及用于上下文相关特征调制的门控反馈。FeedbackTrack保留原始跟踪流程,仅添加固定大小的单帧缓存。在SPMTrack和ARTrackV2上,FeedbackTrack在LaSOT和GOT-10k的5种骨干配置上均实现性能提升,使用SPMTrack-G时达到83.4 AO和79.1 AUC,且参数增加不足1%。对照实验显示,跨帧反馈比同帧调制高出1.8-3.2 AO点,表明性能提升主要来自历史信息的循环利用。进一步分析揭示了学习到的反馈强度存在非均匀的深度依赖组织,凸显循环反馈对Transformer跟踪的有效性。

英文摘要

Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction. Existing temporal mechanisms typically update templates, prompts, queries, or prediction states, while intermediate representations are rarely reused to modulate corresponding processing stages. We propose \textbf{FeedbackTrack}, a visual-cortex-inspired framework that introduces sparse, group-level layer-aligned cross-frame feedback into pretrained Transformer trackers. Previous-frame intermediate states are detached, cached, and returned to corresponding Transformer groups in the current frame through two lightweight pathways: Query Feedback for token-level query modulation and Gate Feedback for context-dependent feature modulation. FeedbackTrack preserves the original tracking pipeline with only a fixed-size one-frame cache. Across SPMTrack and ARTrackV2, FeedbackTrack consistently improves five backbone configurations on LaSOT and GOT-10k, achieving 83.4 AO and 79.1 AUC with SPMTrack-G while adding less than 1\% parameters. Controlled comparisons show that cross-frame feedback outperforms same-frame modulation by 1.8--3.2 AO points, demonstrating that the gains mainly come from recurrent historical information. Further analysis reveals a non-uniform depth-dependent organization of learned feedback strengths, highlighting the effectiveness of recurrent feedback for Transformer tracking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑