arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于泛化到未见机器人操作的动力学感知元模仿

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Zhenduo Shang, Xiyao Liu, Bohan Li, Xudong Wang, Teng Ren, Lianqing Liu, Zhi Han

arXiv 2607.15880首次发表:更新:

发表机构

State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Shenyang University of Technology(中国科学院沈阳自动化研究所机器人学国家重点实验室; 中国科学院大学; 沈阳工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机器人模仿学习的数据稀缺与环境泛化问题,提出动力学感知元模仿(DAMI)框架,通过整合元学习、引入视觉运动轨迹模块等方法,有效捕捉任务动力学,在模拟和真实场景实验中表现优于现有基线。

AI 中文摘要

模仿学习旨在让机器人从大量观察和演示中学习技能,存在数据稀缺和环境泛化问题。现有方法主要关注领域内任务模仿,难以泛化到未见任务。为此提出动力学感知元模仿(DAMI)框架,通过整合元学习构建共享技能空间,引入视觉运动轨迹模块捕捉任务潜在空间复杂时空动力学,并提出无配对统一任务模块融合非结构化多模态观察,还集成任务条件特征调制机制。实验表明该框架在直接推理和少样本微调方面优于现有基线。

英文摘要

Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the \textbf{D}ynamics-\textbf{A}ware \textbf{M}eta-\textbf{I}mitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑