arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单一策略,多种具现体:面向异构具身操纵的以相机为中心的统一动作几何预训练

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Xiaomi Embodied Intelligence Team, University of Macau, :, Shaoqing Xu, Fang Li, Guozhi Zhan, Zhixiang Duan, Yuhan Wang, Yuechen Luo, Shengyin Jiang, Hanbing Li, Zhiying Du, Longlong Wang, Longmei Jiang, Weixiang Liang, Ying Gong, Yong Pan, Ziping Zhao, Zhiyuan Chen, Yangwei You, Kun Ma, Qinyuan Liu, Hangjun Ye, Zhi-xin Yang

arXiv 2608.26058首次发表:更新:

发表机构

Xiaomi Embodied Intelligence Team; University of Macau(小米具身智能团队; 澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对异构具身操纵数据的问题,提出UCAG-P预训练方法,通过以相机为中心的统一动作公式对齐数据,在多个基准任务上取得优异性能,无需特定微调。

AI 中文摘要

通用视觉-语言-动作(VLA)策略的扩展严重受限于具身数据的固有异构性,这些数据涵盖不同的机器人形态、相机配置和低级动作空间。现有范式通常通过显式动作重定向、人到机器人的视频合成或特定数据集的适配分支来解决这种不匹配,从根本上阻碍了统一策略的联合学习。我们引入UCAG-P,一种以相机为中心的统一动作公式,其将异构具身数据集在结构上对齐到共享的几何动作空间。UCAG-P并非将机器人特定命令视为共享策略目标,而是通过图像和相机帧坐标中相机可观测的锚定运动来表示操纵,将机械臂、人形机器人和人手视为共同动作模式的不同具现体。几何条件动作转换器将预测的运动与目标具现体的运动学结合以生成可执行控制。所得的解耦架构允许共享VLA策略学习可迁移的操纵几何,同时保留具现体特定的可控性。UCAG-P在4.03K小时的机器人和模拟数据以及2.34K小时的人类演示上进行训练,单个检查点在LIBERO上达到98.3%,在RoboTwin的Easy和Hard任务上分别达到88.7%和89.2%,在LIBERO-Plus上零样本达到82.0%,在RoboCasa GR-1上达到62.0%,无需针对基准进行特定微调。

英文摘要

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.

CommentsTechnical Report,Project page: https://public-bots.github.io/UCAG-P

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑