arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30595cs.CVcs.AI

动作强制:通过恢复潜在自我运动基座在无监督视频上训练世界模型

Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出动作强制方法,通过主成分分析从无标注视频恢复自我运动基座以生成接地动作监督,训练可控世界模型,并引入无参考评估指标,验证模型能学习反向、缩放及组合控制。

中文摘要 AI 辅助

训练可控世界模型需要同步的动作标注,而这些数据集仍然难以获取。现有方法依赖于带有校准传感器的仪器平台、昂贵的人工标注,或缺乏接地性的潜在动作模型。我们转而通过恢复(无需训练)数据派生的自我运动基座,将普通无标注视频转化为动作监督的训练数据。我们跨帧跟踪像素位移,并利用自我运动引起的周期性相干结构直接获得接地的控制信号。使用像主成分分析这样简单的方法即可实现这一点,我们发现主要成分提供有符号、可缩放且可组合的油门-偏航控制,尽管该方法只能恢复数据中表示的运动轴。为防止高容量视频扩散Transformer利用像素级监督,一个在线潜在评论家蒸馏一个冻结的解码器-跟踪器-主成分分析(PCA)教师,而不通过解码器或跟踪器进行反向传播。最后,我们批评使用视频生成指标来评估世界模型,并介绍一个替代性、无参考评估方法的示例。我们测量可控性、合理性、幻化(凭空创造物体)和几何完整性,揭示传统视频指标遗漏的失败。我们表明大多数基线遵循熟悉的动作方向,但难以反向或保持静止。我们的模型在保持组合控制和生成质量的同时处理两者。尽管反向动作仅占我们训练数据的不到1%,我们发现模型学习反向、线性缩放其响应,并将油门与转向组合,所有这些仅通过接地动作空间学习。

英文摘要

Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the method can recover only motion axes represented in the data. To prevent a high-capacity video DiT from exploiting pixel-level supervision, an online latent critic distils a frozen decoder--tracker--PCA (Principal Components Analysis) teacher without backpropagating through the decoder or tracker. Finally we critique the use of video generation metrics to evaluate WMs and introduce an example of an alternative, reference-free evaluation method. We measure \textit{controllability}, \textit{plausibility}, \textit{conjuring} (creating objects out of thin air) and \textit{geometric integrity}, revealing failures that conventional video metrics miss. We show that most baselines follow familiar action directions but struggle to reverse or remain stationary. Our model handles both while retaining compositional control and generation quality. Despite backwards actions being less than $1\%$ of our training data, we find that the model learns to reverse, scale its response linearly, and compose throttle with steering, all simply by learning through a grounded action space.

↑