CL4D:用于动态场景视觉-语言推理的对比语言-4D预训练
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
- University of Moratuwa(莫拉图瓦大学)
- Singapore-MIT Alliance for Research & Technology (SMART) Centre(新加坡-麻省理工研究与技术联盟(SMART)中心)
- Agency for Science, Technology and Research (A*STAR)(新加坡科学、技术与研究局(A*STAR))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出CL4D(首个4D视觉编码器)及基于其的4DVLM,在自建DynAction4D数据集上训练,CL4D性能较现有方法提升约16.75%,4DVLM优于Gemini、GPT-5等前沿视频VLM。
AI中文摘要:
4D理解与推理是具身智能智能体在动态物理环境中运行的基础能力。然而,现有视觉编码器大多局限于静态2D图像或无时间建模的3D点云,或缺乏精确几何深度推理的2D视频。因此,当前方法无法联合捕捉动态场景中的空间结构与运动演化。我们提出CL4D,首个直接在动态点云上运行的基础4D视觉编码器,采用对比学习目标训练,以对齐时空几何表征与自然语言描述。通过学习文本与4D场景动力学间的共享嵌入空间,CL4D可在动态环境中实现零样本运动到文本、文本到运动的检索,并作为基础4D视觉编码器服务于下游4D视觉-语言任务。基于该编码器,我们引入4DVLM,一种基于动态几何表征进行语言生成的4D视觉-语言模型(VLM),它是首个不依赖2D图像、2D视频或静态3D点云、直接在4D点云上运行的VLM。我们在新构建的DynAction4D数据集上训练CL4D,再训练4DVLM,该数据集涵盖不同物体交互与场景环境中的多样人类运动。在多个4D人类动作基准上的大量实验表明,CL4D实现了最先进的性能,较现有方法提升约16.75%;此外,即使前沿视频VLM如Gemini和GPT-5获得与4DVLM所用4D点云对应场景的RGB视频序列,4DVLM仍优于这些模型。
英文摘要:
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds without temporal modeling, or to 2D videos that lack accurate geometric depth reasoning. Consequently, current approaches fail to jointly capture spatial structure and motion evolution in dynamic scenes. We present CL4D, the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions. By learning a shared embedding space between text and 4D scene dynamics, CL4D enables zero-shot motion-to-text and text-to-motion retrieval in dynamic environments and serves as a foundational 4D vision encoder for downstream 4D vision-language tasks. Building on this encoder, we introduce 4DVLM, a 4D vision-language model that conditions language generation on dynamic geometric representations. 4DVLM is the first VLM designed to operate directly on 4D point clouds without relying on 2D images, 2D videos, or static 3D point clouds. We train CL4D and subsequently 4DVLM on a newly constructed dataset termed DynAction4D capturing diverse human motions across varying object interactions and scene environments. Extensive experiments across multiple 4D human action benchmarks demonstrate that CL4D achieves state-of-the-art performance, with improvements of approximately ~16.75% over prior methods. Furthermore, 4DVLM outperforms frontier video VLMs such as Gemini and GPT-5 even when these models are provided with RGB video sequences corresponding to the same scenes represented as 4D point clouds for 4DVLM.