arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时序强制:面向视觉-语言-动作模型的四维表示对齐

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

Xingyu Ding, Yuzhong Zhao, Chunhai Zhao, Yinghuan Shi, Chaoyang Zhao, Yifan Zhang

arXiv 2608.30643首次发表:更新:

发表机构

Nanjing University; Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences(南京大学; 中国科学院自动化研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言-动作模型缺乏时序信息的问题,提出 Temporal Forcing 方法,通过历史通路与时序一致的四维几何特征对齐,在 LIBERO 等任务上取得性能提升。

AI 中文摘要

近期的视觉-语言-动作(VLA)方法通过将其表示与三维场景几何对齐来提升操作性能,但这些方法常因缺乏时序信息而难以处理长 horizon 操作及视觉相似状态间的观测混淆问题:三维场景几何仅能捕获当前状态,无法反映其随时间的演变。为解决此问题,我们提出 Temporal Forcing,一种面向 VLA 模型的四维表示对齐方法。具体而言,我们首先引入历史通路,使基础 VLA 模型能将观测历史总结为具有时序感知的隐表示;随后将这些隐表示与预训练四维基础模型提取的几何特征对齐,该模型通过时序一致的几何表示捕获演变的三维世界,从而实现对动态环境的更深入理解。Temporal Forcing 在 LIBERO 上达到 98.8%,较其基础模型提升 2.2 个百分点;在物理隐藏放置任务上,其全任务成功率从 20.0%提升至 43.3%,代码将公开可用。

英文摘要

Long-horizon robotic manipulation requires vision-language-action (VLA) models to track scene states and their evolution beyond the current observation. However, simply conditioning policies on observation history does not guarantee that the history is effectively utilized: action supervision constrains what the policy should do, but only indirectly constrains what its history representations should retain. To resolve this, we present Temporal Forcing, a 4D representation alignment framework that explicitly supervises latent temporal states and their transitions. Specifically, we first introduce a history pathway that compresses past observations into compact latent tokens. We then align these tokens and current-frame features with geometric targets from a pretrained 4D foundation model, providing direct supervision at both the state and transition levels. The 4D foundation model and alignment heads are used only for training-time supervision. Temporal Forcing improves average success from 96.6% to 98.8% on LIBERO, with the largest gain on LIBERO-Long (93.8% to 97.2%), and from 53.5% to 62.8% across twelve RoboTwin 2.0 tasks. Furthermore, Temporal Forcing increases full-task success from 20.0% to 43.3% on a physical multi-stage hidden-placement task. Controlled experiments show that 4D representation alignment is crucial for making observation history beneficial to the model. Code will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑