arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UNITAS:用于具身操作的原生三维世界动作模型

UNITAS: A 3D-Native World Action Model for Embodied Manipulation

Ruixiang Wang, Yongyi Su, Wenlve Zhou, Bo Yue, Hengyan Liu, Dekun Lu, Yuxin Tian, Yihan Fang, Zerui Wu, Xing Hu, Jietao Chen, Yong Guo, Ziyan He, Junbin Yuan, Guiliang Liu, Xiaofen Xing, Kui Jia

arXiv 2610.12099首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; DexForce Technology Co., Ltd.; Foshan University; Sun Yat-sen University; South China University of Technology(香港中文大学(深圳); 德方科技有限公司; 佛山大学; 中山大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个原生三维世界动作模型UNITAS,通过统一三维空间内的观测、动作与场景动力学,在RoboTwin、LIBERO等基准上实现了更优的动作条件场景预测与具身操作性能。

AI 中文摘要

世界动作模型(WAMs)旨在回答一个耦合的物理问题:给定一项任务指令,机器人应执行何种运动,且该运动将如何改变周围环境?大多数现有WAMs基于预训练视频生成器构建,通过图像或视觉潜变量表示世界演化。然而,机器人交互发生在度量三维空间中,而图像是依赖视角的投影,其像素距离无法直接编码物理距离。我们提出UNITAS,据我们所知,这是首个原生三维世界动作模型,在每次交互中于共享度量三维框架内统一观测、动作与场景动力学,采用适用于机器人实体和人手的通用表示。动作流将人手和机器人夹爪表示为三维点轨迹,而场景流描述受这些轨迹条件约束的场景点位移。与世界对齐的三维位置嵌入将视觉标记(无论是否有深度输入)锚定,物理时间轨迹分词器将每个点轨迹编码为一个锚定在其当前三维位置的标记。该接口支持直接动作执行和动作条件下的场景预测。UNITAS拥有17亿参数,在RoboTwin数据集上的动作条件场景预测性能优于所有对比方法,位移误差较PointWorld低达49%,且实现了具身操作的最优性能,包括在LIBERO任务上达到99.8%的成功率,在真实世界任务中的平均成功率为85%。代码可在指定URL获取。

英文摘要

World action models (WAMs) aim to answer a coupled physical question: given a task instruction, what motion should the robot execute, and how will that motion change the surrounding world? Most existing WAMs build on pretrained video generators and represent world evolution through images or visual latents. Robotic interaction, however, takes place in metric three-dimensional space, while images are view-dependent projections whose pixel distances do not directly encode physical distances. We introduce UNITAS, to our knowledge the first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot embodiments and human hands. Action flow represents human hands and robot grippers as 3D point trajectories, while scene flow describes scene-point displacements conditioned on these trajectories. World-aligned 3D positional embeddings ground visual tokens with or without depth input, and a physical-time trajectory tokenizer encodes each point trajectory as one token anchored at its current 3D position. This interface supports both direct action execution and action-conditioned scene prediction. With 1.7B parameters, UNITAS achieves the best action-conditioned scene prediction on RoboTwin among the compared methods, with up to 49% lower displacement errors than PointWorld, and state-of-the-art manipulation success, including 99.8% on LIBERO and an average of 85% across real-world tasks. The code is available at https://github.com/DexForce/UNITAS.

Comments27 pages, 11 figures and 13 tables. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑