arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37348cs.RO

DROM:一种语言引导的多技能机器人操作扩散框架

DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation

Vincenzo Pomponi, Rocco Felici, Paolo Franceschi, Stefano Baraldo, Oliver Avram, Loris Roveda, Luca Maria Gambardella, Anna Valente

首次发表
浏览论文内容

中文总结 AI 辅助

DROM提出语言引导的扩散框架,利用DMPs扩充示范并扩展MPD,实现多技能操作与长时程任务,在真实机器人和模拟中优于基线,仅需少量示范。

中文摘要 AI 辅助

从有限的示范中学习多样、长时程任务的鲁棒操作策略仍然是机器人学中的一个基本挑战。我们提出了DROM,一种语言引导的扩散框架,使机器人能够在单个生成策略中学习、表示和组合多种操作技能。DROM利用动态运动基元(DMPs)将少量专家示范扩充为表现力丰富的多技能数据集,大幅减少数据收集,同时提高超出示范工作空间的空间泛化能力。基于运动规划扩散(MPD),我们扩展了扩散架构,通过交叉注意力支持语言条件下的多技能轨迹生成,使单个模型能够为多样化的操作基元生成技能一致的轨迹,包括传统运动规划或硬编码控制器难以设计的对方向敏感的行为。对于长时程操作,大型语言模型将高层操作员请求分解为可执行的技能序列,实现自然语言交互和自主任务执行。我们在Franka Emika Panda机器人、FANUC CRX25ia机器人以及MuJoCo模拟环境中,在广泛的操作任务上验证了DROM。实验结果表明,DROM优于运动规划扩散和行为克隆基线,实现了鲁棒的多技能泛化,并仅使用有限数量的人类示范,从自然语言指令中组合所学技能以可靠地执行长时程操作任务。数据集、模拟环境等更多信息见此https URL。

英文摘要

Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffusion (MPD), we extend the diffusion architecture to support language-conditioned multi-skill trajectory generation through cross-attention, allowing a single model to generate skill-consistent motions for a diverse set of manipulation primitives, including orientation-sensitive behaviors that are difficult to design using conventional motion planning or hard-coded controllers. For long-horizon manipulation, a large language model decomposes high-level operator requests into executable sequences of skills, enabling natural language interaction and autonomous task execution. We validate DROM on a Franka Emika Panda robot, a FANUC CRX25ia robot, and in MuJoCo simulation across a wide range of manipulation tasks. Experimental results demonstrate that DROM outperforms Motion Planning Diffusion and Behavior Cloning baselines, achieves robust multi-skill generalization, and composes learned skills to reliably execute long-horizon manipulation tasks from natural language instructions using only a limited number of human demonstrations. Datasets, simulation environments, and more at https://github.com/automation-robotics-machines/drom.

发表机构

  • University of Applied Science and Arts of Southern Switzerland (SUPSI)(瑞士南部应用科学与艺术大学)
  • Institute of Systems and Technologies for Sustainable Production (ISTePS)(可持续生产系统与技术研究所)
  • Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA)(达勒莫尔人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑