StrucPhysVideo:从结构化描述和机器人动作学习物理动力学
StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions
- Awomo-WM Team(Awomo-WM团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
StrucPhysVideo通过结构化物理数据流程和课程学习,提出视频世界模型系列,在Physics-IQ上达到SOTA,并扩展至动作驱动的交互预测。
AI中文摘要:
对物理动力学(包括物体如何移动、相互作用和改变状态)进行建模,是具身人工智能视频世界模型的核心。我们提出了StrucPhysVideo,一个视频世界模型系列,它将物理聚焦的数据整理与语言和动作条件下的场景演化预测相结合。我们的数据流程结合了运动感知的视频分割、质量和内容过滤、物理相关性验证,以及物体、材料和时域局部化交互的结构化注释。通过将相机运动与物体行为分离,并明确描述接触、变形和状态转换,该流程提供了基于可观察物理事件的监督。在这些数据的基础上,我们引入了StrucPhysVideo-TI2V,一个稀疏专家混合(MoE)文本-图像到视频模型,通过课程学习进行训练,逐步强调物理动力学,同时保留通用领域的视频数据。StrucPhysVideo-TI2V在Physics-IQ Verified上取得了最先进的性能,得分45.5%,比Cosmos3-Super-Image2Video高出2.8个百分点。跨骨干网络的描述消融实验进一步证明了物理聚焦监督的有效性。我们进一步将StrucPhysVideo-TI2V扩展到StrucPhysVideo-IA2V,一个交互式图像-动作到视频世界模型,它从机器人末端执行器命令预测视觉结果。动作条件、因果自回归生成和少步蒸馏使得仅用四步去噪即可实现增量式机器人推演。总之,StrucPhysVideo将物理动力学建模从图像和语言条件下的视频预测推进到动作驱动的交互。
英文摘要:
Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucPhysVideo, a family of video world models that bridges physics-focused data curation with language- and action-conditioned prediction of scene evolution. Our data pipeline combines motion-aware video segmentation, quality and content filtering, and physical relevance verification with structured annotations of objects, materials, and temporally localized interactions. By disentangling camera motion from object behavior and explicitly describing contact, deformation, and state transitions, the pipeline provides supervision grounded in observable physical events. Building on these data, we introduce StrucPhysVideo-TI2V, a sparse Mixture-of-Experts (MoE) text-image-to-video model trained with a curriculum that progressively emphasizes physical dynamics while retaining general-domain video data. StrucPhysVideo-TI2V achieves state-of-the-art performance on Physics-IQ Verified, scoring 45.5% and outperforming Cosmos3-Super-Image2Video by 2.8 percentage points. Caption ablations across backbones further demonstrate the effectiveness of physics-focused supervision. We further extend StrucPhysVideo-TI2V to StrucPhysVideo-IA2V, an interactive image-action-to-video world model that predicts visual outcomes from robot end-effector commands. Action conditioning, causal autoregressive generation, and few-step distillation enable incremental robot rollouts with only four denoising steps. Together, StrucPhysVideo advances physical dynamics modeling from image- and language-conditioned video prediction toward action-driven interaction.