arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14028cs.ROcs.AI

AdvDex:通过关节对齐动作与对抗学习从人类演示中学习灵巧操作

AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

发表机构浙江大学 · 上海创新研究院 · 复旦大学
另 2 家 · 查看机构详情
  • Zhejiang University(浙江大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Fudan University(复旦大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Paxini Tech(帕西尼科技)

机构由 AI 辅助整理,请以论文原文为准。

Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang, Tong He, Zhengxue Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

AdvDex是一种统一视觉-语言-动作框架,通过OmniShare数据集、JAAS动作空间与领域对抗学习,实现从人类和机器人演示中学习灵巧操作,提升跨实体泛化与技能迁移能力。

中文摘要 AI 辅助

灵巧操作是具身智能的基础能力,但规模化实现十分困难,因为机器人演示的采集成本高昂,且不同实体的动作空间存在差异。在异构数据上训练的策略还可能将任务相关视觉线索与实体特定外观纠缠,限制跨实体泛化能力。本文提出AdvDex,一种用于从人类和机器人演示中学习灵巧操作的统一视觉-语言-动作框架。首先,引入OmniShare——大规模多模态人类操作演示数据集,提供高质量运动学监督与触觉测量,同时减少对机器人遥操作的依赖。其次,提出关节对齐动作空间(JAAS),一种规范动作表示,包含SE(3)腕部位姿与15个手指关节,从而在功能上对齐人类手、灵巧机器人手与平行夹爪。最后,采用领域对抗学习减少学习到的视觉表示中的实体特定信息。在手部动作预测与现实世界灵巧操作上的实验显示,其较基线方法有持续提升,具备有效的零样本人机技能迁移、对未见过的物体与环境的泛化能力,以及数据高效的少样本适配。

英文摘要

Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can also entangle task-relevant visual cues with embodiment-specific appearance, limiting cross-embodiment generalization. We present AdvDex, a unified Vision-Language-Action framework for learning dexterous manipulation from human and robot demonstrations. First, we introduce OmniShare, a large-scale multimodal dataset of human manipulation demonstrations that provides high-quality kinematic supervision and tactile measurements while reducing reliance on robot teleoperation. Second, we propose the Joint-Aligned Action Space (JAAS), a canonical action representation comprising an $\mathrm{SE}(3)$ wrist pose and 15 finger joints, thereby functionally aligning human hands, dexterous robot hands, and parallel grippers. Finally, we use domain-adversarial learning to reduce embodiment-specific information in the learned visual representation. Experiments on hand-action prediction and real-world dexterous manipulation show consistent improvements over baselines, effective zero-shot human-to-robot skill transfer, generalization to unseen objects and environments, and data-efficient few-shot adaptation.

↑