arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16695cs.CV

MAETrack:释放预训练几何先验在3D单目标跟踪中的潜力

MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking

Sifan Zhou, Qiwei Wang, Linyue Tan, Ziyu Liu, Ziyu Zhao, Xiaobo Lu

首次发表
浏览论文内容

中文总结 AI 辅助

针对3D单目标跟踪中预训练MAE表示迁移不佳的问题,提出MAETrack框架,通过层选择性初始化和几何残差门控,在低开销下提升跟踪性能。

中文摘要 AI 辅助

大规模预训练已经变革了2D视觉中的表示学习,但其向3D单目标跟踪(SOT)的可迁移性仍未得到充分理解。直接微调自监督3D编码器(如掩码自编码器MAE)往往导致次优的适应效果,因为重建目标与跟踪的空间-时间匹配需求并未完全对齐。在本文中,我们观察到这一困难可以解释为逐层迁移不匹配:浅层倾向于保留可迁移的几何线索,而深层则逐渐专门化于重建 pretext 任务,不太适合下游跟踪。基于这一观察,我们提出了MAETrack,一个轻量级适应框架,用于将预训练的MAE表示迁移到3D SOT。MAETrack包括层选择性初始化(LSI),该初始化仅从预训练权重初始化跟踪骨干的浅层阶段,而重新初始化更深阶段;以及几何残差门控(GRG),该门控在模板-搜索融合之前,通过残差空间调制增强搜索BEV特征中结构显著的区域。在标准3D SOT基准上的大量实验表明,MAETrack在有限的计算开销下持续优于普通微调基线。更广泛地,我们的结果表明,从3D重建预训练到3D跟踪的有效迁移不仅仅是部分微调的问题,而是依赖于一种面向跟踪的迁移原则,该原则保留浅层几何信息,同时使深层表示适应下游目标。

英文摘要

Large-scale pre-training has transformed representation learning in 2D vision, yet its transferability to 3D single object tracking (SOT) remains insufficiently understood. Directly fine-tuning self-supervised 3D encoders, such as masked autoencoders (MAE), often leads to sub-optimal adaptation because the reconstruction objective is not fully aligned with the spatial-temporal matching requirements of tracking. In this paper, we observe that this difficulty can be interpreted as a layer-wise transfer mismatch: shallow layers tend to preserve transferable geometric cues, while deeper layers become increasingly specialized to the reconstruction pretext task and are less suitable for downstream tracking. Based on this observation, we propose MAETrack, a lightweight adaptation framework for transferring pre-training MAE representations to 3D SOT. MAETrack includes Layer-Selective Initialization (LSI), which initializes only the shallow stages of the tracking backbone from pre-trained weights while re-initializing deeper stages, and Geometric Residual Gating (GRG), which reinforces structurally salient regions in the search BEV features before template-search fusion through residual spatial modulation. Extensive experiments on standard 3D SOT benchmarks show that MAETrack consistently improves upon vanilla fine-tuning baselines with limited computational overhead. More broadly, our results suggest that effective transfer from 3D reconstruction pre-training to 3D tracking is not merely a matter of partial fine-tuning, but depends on a tracking-oriented transfer principle that preserves shallow geometry while adapting deeper representations to the downstream objective.

发表机构

  • Southeast University(东南大学)
  • Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳))
  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑