发表机构
National Tsing Hua University; Comfy Org; Research realiz.ai; Karolinska Institutet; Stockholm University; NVIDIA(清华大学; Comfy组织; realiz.ai研究院; 卡罗林斯卡学院; 斯德哥尔摩大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频自监督学习中方法比较混杂的问题,提出 TT-VidT,通过解耦时间轴并结合 Diff Compression 训练,在多个数据集微调中显著领先,同时大幅降低计算开销。
AI 中文摘要
视频自监督学习中的比较往往评估完整的训练方案,而非隔离方法本身:架构、目标、数据暴露、调度、规模和解码器容量都可能同时变化。这使得难以识别哪些选择能产生以运动优先的表征,其增益集中在帧间变化上,同时保留有用的外观信息。我们通过一项匹配的 $4 \ imes 6 = 24$ 架构-目标研究来解决这一问题,该研究在约 1.7M OpenVid 和 Moments-in-Time v2 片段上,以约 1.7 亿至 1.9 亿编码器规模进行 8 个 epoch 的训练,并提出了 TT-VidT。TT-VidT 将 DINOv3 初始化的 ViT-B/16 逐帧空间路径与紧凑的时间转移层相结合,通过 Diff Compression 训练,从首帧外观锚点和特定帧的运动标记重建目标帧。扫描结果表明,TT3D 与 Diff Compression 结合(而非任一单独组件)进入了最强的运动敏感区域,且解码器消融实验倾向于紧凑的视频预训练解码器。在最终比较中,TT-VidT 同时引领了 Jester、Something-Something V2、ARID 和 Diving48 的微调,相比最强的非 TT 行提升了 54% 至 121%,同时相比 DisMo 减少了 48% 的编码器 FLOPs,相比 VideoMAE 或 V-JEPA2 减少了 55%。HMDB51、IARD 和 EPIC-Kitchens 界定了该主张的适用范围。
英文摘要
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
CommentsAccepted by NeurIPS 2026 main track, Project page: https://kohakublueleaf.github.io/TTVidT/