arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09514cs.CV

STRIKE:学习视觉状态转换以进行物理世界建模

STRIKE: Learning Visual State Transitions for Physical World Modeling

  • Applied Intuition(应用直觉公司)
  • University of Southern California(南加州大学)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Wenbin Teng, Tianshuo Xu, Depu Meng, Yuelei Li, Quentin Herau, Yihan Hu, Yajie Zhao, Wei Zhan

AI总结:

STRIKE通过分离视觉状态转换学习与视频生成,利用事件对齐监督和递归预测,在多个基准上提升物理一致性与操作视频保真度。

AI中文摘要:

物理世界建模需要预测交互如何改变场景,而不仅仅是生成连贯的运动。我们提出STRIKE,一个将视觉状态转换学习与密集视频生成分离的框架。我们通过从训练视频中提取观察到的状态,并将其与转换描述和时间偏移配对,构建事件对齐的监督。一个基于图像的转换模型学习从当前图像、局部转换规范和经过的时间预测下一个场景配置。在推理时,一个预训练的视觉-语言规划器预测时间转换规范,并且学习到的转换模型的递归应用产生一系列未来视觉状态。一个单独训练的动力学模型然后基于这些状态及其时间位置生成完整的展开。在Physics-IQ Verified、PhyGenBench、Pisa-Experiments和RoboTwin2.0上的实验显示,STRIKE在物理一致性和操作视频保真度的基准度量上相对于相应的视频骨干基线有所改进。这些结果支持学习到的视觉状态转换作为物理世界建模的有效中间表示。

英文摘要:

Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

↑