发表机构
DAMO Academy, Alibaba Group; Hong Kong Embodied AI Lab; CUHK; Hupan Lab(达摩院,阿里巴巴集团; 香港具身人工智能实验室; 香港中文大学; 湖畔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对开放世界机器人操作,提出RynnWorld-4D生成模型,通过RGB-DF捕捉4D动态,其具三分支架构,还整理数据集并提出RynnWorld-4D-Policy,实验证明该模型在时空预测及实际操作任务中表现出色。
AI 中文摘要
开放世界中的机器人操作不仅需要识别场景外观,还需预测其3D结构在交互中的移动。我们认为同步的RGB、深度和光流(RGB-DF)提供了一种基于物理的表示,能捕捉场景的潜在4D动态。基于此,我们引入RynnWorld-4D,一个在统一扩散过程中从单个RGB-D图像和语言指令共同生成未来RGB帧、深度图和光流的生成模型。它具有三分支架构,集成跨模态注意力和逐帧3D RoPE。我们还整理了Rynn4DDataset 1.0数据集,并提出RynnWorld-4D-Policy。实验表明,RynnWorld-4D能产生时空连贯的4D预测,RynnWorld-4D-Policy在实际灵巧双手操作任务中达到了先进性能。
英文摘要
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow (RGB-DF) provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to low-level end-effector actions demanded by robotic systems, narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.
CommentsProject Page: https://alibaba-damo-academy.github.io/RynnWorld-4D.github.io, Github: https://github.com/alibaba-damo-academy/RynnWorld-4D