发表机构
Beijing Humanoid Robot Innovation Center; Beijing Institute of Technology; Harbin Institute of Technology; Shenzhen; The University of Hong Kong; China University of Mining & Technology, Beijing(北京人形机器人创新中心; 北京理工大学; 哈尔滨工业大学; 深圳; 香港大学; 中国矿业大学(北京))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Dream4ACT提出共享视觉动作接口(动作视图)统一多具身动作表示,结合掩码流匹配实现联合视频-动作建模,无需训练恢复可执行动作,在RoboTwin 2.0和TriWorldBench上取得领先性能。
AI 中文摘要
视频生成模型(VGM)为具身观测-动作建模提供了强大的时空先验。然而,关节空间的动作向量缺乏显式的图像空间结构,且在不同具身之间维度和语义各异,这使得直接利用VGM丰富的时空先验变得困难。末端执行器可视化提供了一种替代方案,但无法指定机器人执行所需的完整关节配置。我们提出了Dream4ACT,一个为跨具身联合视频-动作建模而构建的世界模型。为了统一不同具身之间的动作表示,我们引入了一种共享的视觉动作接口,称为动作视图,它利用基于URDF的正向运动学,从四个预设的虚拟相机渲染目标关节配置。这种共享的视觉表示保留了具身特有的关节几何结构,同时允许观测和动作序列共享一个视频自编码器和扩散变换器。通过掩码流匹配,我们的模型通过改变哪些未来序列被破坏,在单个联合训练的模型中支持正向动力学、逆向动力学以及联合观测-动作生成。为了从预测的动作视图中恢复可执行的动作序列,我们提出了一种无需训练、受URDF约束的多视图恢复机制,无需学习具身特有的解码器。Dream4ACT在RoboTwin~2.0上实现了88.98%的平均成功率,在TriWorldBench上获得了65.66的总体得分,通过视觉动作接口支持有效的闭环操作和具有竞争力的动作条件多视图预测。
英文摘要
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.