DemoBot: 从第三人称人类视频高效学习双臂操作与灵巧手
DemoBot: Efficient Learning of Bimanual Manipulation with Dexterous Hands From Third-Person Human Videos
浏览论文内容
中文总结 AI 辅助
DemoBot通过强化学习框架从人类视频中高效学习双臂操作技能,解决长周期任务的挑战。
中文摘要 AI 辅助
本工作提出了DemoBot,一种学习框架,使双臂、多指机械系统能够从单个未标注的RGB-D视频演示中获取复杂的操作技能。该方法从原始视频数据中提取双臂和物体的结构化运动轨迹。这些轨迹作为新型强化学习(RL)管道的运动先验,通过接触丰富的交互来学习优化它们,从而避免了从头开始学习的需要。为了解决学习长周期操作技能的挑战,我们引入了:(1)基于时间段的RL以强制当前状态与演示的时序对齐;(2)成功门控重置策略以平衡已获得技能的优化与后续任务阶段的探索;(3)事件驱动的奖励课程配以自适应阈值以指导RL学习高精度操作。新颖的视频处理和RL框架成功实现了长周期同步和异步双臂装配任务,提供了一种可扩展的方法,直接从人类视频中获取技能。
英文摘要
This work presents DemoBot, a learning framework that enables a dual-arm, multi-finger robotic system to acquire complex manipulation skills from a single unannotated RGB-D video demonstration. The method extracts structured motion trajectories of both hands and objects from raw video data. These trajectories serve as motion priors for a novel reinforcement learning (RL) pipeline that learns to refine them through contact-rich interactions, thereby eliminating the need to learn from scratch. To address the challenge of learning long-horizon manipulation skills, we introduce: (1) Temporal-segment based RL to enforce temporal alignment of the current state with demonstrations; (2) Success-Gated Reset strategy to balance the refinement of readily acquired skills and the exploration of subsequent task stages; and (3) Event-Driven Reward curriculum with adaptive thresholding to guide the RL learning of high-precision manipulation. The novel video processing and RL framework successfully achieved long-horizon synchronous and asynchronous bimanual assembly tasks, offering a scalable approach for direct skill acquisition from human videos.
发表机构
- ByteDance Seed(字节跳动种子)
机构由 AI 辅助整理,请以论文原文为准。