发表机构
Fundamental AI Lab, UTN; VGG, University of Oxford(基础人工智能实验室,UTN; VGG,牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究能否从预训练图像视觉Transformer的冻结特征中恢复结构化运动表示,提出结构化动力学模型(SDM),通过未来特征预测分离动力学来源,结合自监督与弱监督训练,在新评估套件上表现出色,证明预训练图像模型可用于结构化视频动力学表示。
AI 中文摘要
理解视频中的运动是视觉学习的一项基本挑战,因为帧间变化包含相机运动和物体运动这两种动力学来源。在表示学习中,这种分解尚未得到充分探索,部分原因是这些因素在自然视频中紧密耦合且难以单独监督。然而,恢复这种分解对于学习将有意义的物体动力学与相机引起的变化分离的鲁棒运动表示很重要。我们研究是否可以从预训练图像视觉Transformer的冻结特征中恢复这种结构化运动表示。我们提出了结构化动力学模型(SDM),它通过未来特征预测明确地将时间变化的主要来源与残余动力学分开,而不是用单个纠缠的潜在特征或无结构的空间密集过渡令牌来表示视频变化。训练将对真实视频的自监督学习与对合成Kubric数据的场景动力学弱监督相结合。我们在ProbeMotion上评估SDM,这是一个新的评估套件,涵盖具有相机运动、物体运动和组合动力学的合成和真实视频。SDM在使用全局CLS或平均池化特征方面优于主干基线,并且在几个探针上与诸如VGGT等强监督表示相比具有优势,尽管使用的监督要弱得多。这些结果表明,预训练的图像模型可以很容易地重新用于结构化视频动力学表示,为学习和分析潜在视频动力学提供有用的归纳偏差。
英文摘要
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.
Commentspreprint, Project page: https://lukasknobel.github.io/projects/StructuredDynamics