发表机构
Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究动作条件视频世界模型,提出机器人因素分解的世界模型,通过动作实现和机器人渲染将特定机器人因素移出模型,解决深度模糊问题,实验表明其优于基线且能推广,还能从人类演示生成机器人操作视频。
AI 中文摘要
动作条件视频世界模型根据初始观测和动作信号预测未来观测。在机器人技术中,动作通过两个不同过程影响未来观测:先由机器人身体和控制器转化为机器人运动,然后场景通过接触和物体运动做出响应。直接以动作命令为条件要求世界模型学习实现过程本身,而以记录的未来状态为条件会泄露其要预测的交互结果。我们提出机器人因素分解的世界模型,将两个特定于机器人的因素移出世界模型。一是动作实现,将每个命令通过机器人自身控制器和运动学转化为可部署的名义轨迹,避免动作实现学习和未来状态泄露。二是机器人渲染,通过机器人URDF渲染名义轨迹,将机器人几何、运动学和外观从模型中分解出来。为解决深度模糊问题,将末端执行器深度与场景深度配对。实验表明渲染接口优于向量条件基线,并能推广到推理时未见的机器人实例。还证明该模型可通过将手部动作重定向并渲染为机器人几何形状,从人类演示中生成机器人操作视频。
英文摘要
Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two distinct processes: they are first realized into robot motion by the robot body and controller, and the scene then responds through contact and object motion. Conditioning directly on action commands asks the world model to learn the realization process itself, while conditioning on logged future states leaks the interaction outcomes it is meant to predict. We propose robot-factored world models, which move two robot-specific factors outside the world model. First, action realization: each command is rolled through the robot's own controller and kinematics into a deployment-available nominal trajectory, a middle signal that avoids both action-realization learning and future-state leakage. Second, robot rendering: this nominal trajectory is rendered through the robot URDF, factoring the robot's geometry, kinematics, and appearance out of the model and into explicit rendered robot geometry. To resolve depth ambiguity, we pair end-effector depth with scene depth, giving geometric cues for contact and occlusion beyond image-plane overlap. Together, camera-aware static RGB/depth context and rendered robot geometry form a shared visual world-model interface that stays consistent across viewpoints and robot embodiments, so the model sees the action only as visible robot geometry and learns how objects respond to it. Our experiments show that the rendered interface outperforms vector-conditioned baselines and generalizes to unseen robot embodiments at inference. We further demonstrate that our model generates robot manipulation videos from human demonstrations by retargeting and rendering the hand motion as robot geometry.
CommentsProject Page: https://bjkim95.github.io/rofacto/