arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12470cs.ROcs.CV

Dex-One2Many:从单次人类演示中学习灵巧操作

Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration

Jusuk Lee, Sungha Kim, Yeonsoo Park, Jonguk Cheon, Yoonkyo Jung, Yongjun You, H. Jin Kim, Jia-Bin Huang, Furong Huang, Youngseok Jang, Seungjae Lee

首次发表
浏览论文内容

中文总结 AI 辅助

Dex-One2Many是一种实境-仿真-实境框架,从单次人类视频学习灵巧操作策略,通过场景图指导强化学习,在五项任务中已见配置性能超基线6.5%、未见场景超71%,实现高效探索与强泛化。

中文摘要 AI 辅助

虽然从单次人类视频中学习灵巧操作是替代昂贵机器人演示的有前景方案,但许多近期方法主要模仿演示动作,这种严格的动作匹配往往限制了对视频中未出现的初始物体位姿、目标位姿和抓取方式的泛化能力。另一方面,通过强化学习(RL)发现策略可实现广泛泛化,但若无先验指导,其在复杂多阶段任务的高维探索中会遇到困难。为解决这一泛化与探索的耦合挑战,我们提出Dex-One2Many,这是一种从单次人类视频中学习可泛化灵巧操作策略的实境-仿真-实境框架。我们的核心见解是将视频抽象为序列场景图以指导RL,从而在保持广泛泛化性的同时实现高效探索。这些场景图作为生成式约束,用于采样多样化的重置状态,并为每个阶段提供密集奖励。由于场景图约束的是关系而非精确位姿,这些重置状态涵盖了视频之外的物体位姿和抓取方式,而通过密集奖励从这些状态初始化每个阶段,可缩短探索过程并使其具有指导性。Dex-One2Many完全在仿真中训练,可零样本迁移至真实多指机械手。在五项工具使用与操作任务中,Dex-One2Many在已见配置上比基线方法高出6.5%,而其强大的泛化能力使这一差距在未见场景中扩大至71%。

英文摘要

While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.

发表机构

  • Seoul National University(首尔大学)
  • University of Maryland, College Park(马里兰大学帕克分校)
  • All Purpose AI
  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑