arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DexJoCo-X:多手灵巧操作的动作表示基准

DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation

Xiangwei Jiang, Yao Mu, Lixin Duan, Wen Li

arXiv 2610.03278首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Shanghai Jiao Tong University(电子科技大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DexJoCo-X基准通过受控实验证明,跨实体灵巧操作需要动作坐标、预训练和架构的联合设计,其中Being-H0.5以47.7%成功率实现单策略控制七种手型。

AI 中文摘要

随着灵巧手的普及,为每种形态分别收集数据和训练策略变得越来越不切实际。因此,可扩展的跨实体学习需要一种统一的表示,既能捕捉共享的操作结构,又能保留特定形态的控制。手部、任务、数据集和控制接口的差异使得现有研究无法隔离表示、预训练和架构的影响。我们引入了DexJoCo-X,一个用于在七种代表性灵巧手、六个单臂和双臂任务以及2,100个平衡演示中进行受控比较的基准和工具包。DexJoCo-X提供了一个匹配的多手、多任务协议,具有共同的场景、成功标准和执行接口,重新设计的从手套到手部的映射,以及一个自动化流水线,可在随机化场景中扩展经过审查的演示。使用π_{0.5}、Ego-Pi和Being-H0.5,我们检验了共享动作接口是否足以支持多手学习。将π_{0.5}扩展到80维双臂输出导致接近零的成功率。Ego-Pi通过交错预测保留了预训练的动作头,并支持每手多任务学习,但在七手联合训练中仍然无效。相比之下,Being-H0.5结合了跨实体预训练、统一动作空间和实体感知专家,使得一个策略能够控制所有七只手。在该架构内,功能对齐的动作槽实现了47.7%的平均成功率,而原生坐标和DexLatent分别为47.0%和33.1%。这些结果表明,跨实体表示依赖于整个学习系统:动作坐标、预训练和架构必须共同将共享的操作结构与特定实体的控制分离。

英文摘要

As dexterous hands proliferate, collecting data and training policies separately for every morphology becomes increasingly impractical. Scalable cross-embodiment learning therefore requires a unified representation that captures shared manipulation structure while preserving morphology-specific control. Differences in hands, tasks, datasets, and control interfaces prevent existing studies from isolating the effects of representation, pretraining, and architecture. We introduce DexJoCo-X, a benchmark and toolkit for controlled comparison across seven representative dexterous hands, six single-arm and bimanual tasks, and 2,100 balanced demonstrations. DexJoCo-X provides a matched multi-hand, multi-task protocol with common scenes, success criteria, and execution interfaces, redesigned glove-to-hand mappings, and an automated pipeline that expands reviewed demonstrations across randomized scenes. Using $π_{0.5}$, Ego-Pi, and Being-H0.5, we examine whether a shared action interface is sufficient for multi-hand learning. Expanding $π_{0.5}$ to an 80-dimensional bimanual output yields near-zero success. Ego-Pi preserves the pretrained action head through interleaved prediction and supports per-hand multi-task learning, but remains ineffective for seven-hand joint training. By contrast, Being-H0.5 combines cross-embodiment pretraining, a unified action space, and embodiment-aware experts, enabling one policy to control all seven hands. Within this architecture, function-aligned action slots achieve 47.7% mean success, compared with 47.0% for native coordinates and 33.1% for DexLatent. These results show that cross-embodiment representation depends on the entire learning system: action coordinates, pretraining, and architecture must jointly separate shared manipulation structure from embodiment-specific control.

Comments8 pages, 5 figures. Project website: https://darenrenjian.github.io/DexJoCo-X-website/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑