发表机构
DreamX(DreamX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DreamX-Phi 1.0是面向机器人操控的动作条件视频世界模型,通过几何编码、深度分支等优化,在WorldArena 2.0挑战赛获Track1第一、Track2第二,模型与代码将公开。
AI 中文摘要
我们提出DreamX-Phi 1.0,这是一种面向机器人操控的动作条件视频世界模型,给定观测帧、语言指令以及包含末端执行器位姿和夹爪状态的规定动作序列,它能够预测后续的观测结果。然而仅靠逼真性无法保证预测的准确性:一段有说服力的 rollout 仍可能移动错误的机械臂或丢失被操控物体。为确保预测符合每个机械臂的指令路径,我们通过PRoPE风格几何编码将各机械臂的SE(3)变换注入注意力机制,以保留机械臂身份和刚体运动结构。仅动作控制无法完全约束场景几何或小型被操控物体的演变,因此我们添加轻量深度分支用于场景级几何建模,并使用结合冻结V-JEPA教师模型的SAM3掩码,以在抓取过程中保持物体一致性。我们还通过分布匹配蒸馏将多步生成器提炼为少步学生模型,以实现高效部署。截至撰写本文时,\nmodel在WorldArena 2.0挑战赛的Track 1中取得第一名,在Track 2中取得第二名。我们的模型和代码将公开提供。
英文摘要
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.
CommentsCode: https://github.com/AMAP-ML/DreamX-Phi