arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分解与重组:利用从示范中学到的原语和视觉运动策略进行规划

Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations

Yizhou Chen, Hang Xu, Dongjie Yu, Yupu Lu, Tengye Xu, Zeqing Zhang, Wei Zhang, Yi Ren, Ben M. Chen, Jia Pan

arXiv 2607.25397首次发表:更新:

发表机构

The University of Hong Kong; JD.com (JingDong); Nanyang Technological University (NTU); Southern University of Science and Technology (SUSTech); Huawei Technologies; The Chinese University of Hong Kong (CUHK)(香港大学; 京东; 南洋理工大学; 南方科技大学; 华为技术有限公司; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何使机器人灵巧、长视野操作自动化,提出DR-LfD框架,将视觉运动策略集成到TAMP决策系统,基于接触关系分解示范为原子技能,经实验验证该框架在多类任务中性能强大。

AI 中文摘要

成功地使灵巧、长视野的机器人操作自动化需要能够进行高级推理和细粒度执行的框架。传统的任务和运动规划(TAMP)在符号规划方面表现出色,但在接触丰富的操作中往往很脆弱。同时,模仿学习(IL)在具有视觉反馈的操作任务中有效,但在空间泛化和多阶段操作方面能力有限。为了协调它们的互补优势和局限性,我们提出了DR-LfD(从示范中学到的分解和重组技能),这是一个将视觉运动策略无缝集成到TAMP门控决策系统中的框架。基于接触关系,DR-LfD将人类示范分解为原子技能,这些技能被再现为视觉运动策略或以对象为中心的原语。视觉运动策略的启动、终止和约束以与TAMP兼容的形式精心建模和实现,从而能够重组从不同来源学到的技能。DR-LfD将学习问题从一个需要在可能的技能序列上使用指数级示范数据的问题转变为一个示范负担随不同技能类型数量而扩展的问题,每个技能的数据有限。通过在各种场景下进行全面的现实世界和模拟基准测试,我们证明了DR-LfD在涉及多个步骤、未见设置和物理约束的任务上的强大性能。

英文摘要

Successfully automating dexterous, long-horizon robotic manipulation requires frameworks capable of both high-level reasoning and fine-grained execution. Traditional task and motion planning (TAMP), while excellent at symbolic planning, is often brittle in contact-rich operations. Simultaneously, imitation learning (IL), while effective in manipulation tasks with visual feedback, is limited by its low capability in spatial generalization and multi-stage operation. To reconcile their complementary strengths and limitations, we propose DR-LfD (Decomposed and Reorganized Skills Learned from Demonstrations), a framework that seamlessly integrates visuomotor policies into a TAMP-gated decision-making system. Based on contact relationships, DR-LfD decomposes human demonstrations into atomic skills, which are reproduced as visuomotor policies or object-centric primitives. The initiation, termination, and constraints of the visuomotor policies are carefully modeled and implemented in a TAMP-compatible form, enabling reorganization of skills learned from different sources. DR-LfD transforms the learning problem from one requiring exponential demonstration data over possible skill sequences to one whose demonstration burden scales with the number of distinct skill types, with limited data for each skill. Through comprehensive real-world and simulation benchmarking across diverse scenarios, we demonstrate the strong performance of DR-LfD on tasks involving multiple steps, unseen setups, and physical constraints. Project website: https://dr-lfd.github.io/DR-LfD-website.

Comments21 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑