发表机构
ISIR, Sorbonne Université(ISIR,索邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出仅需自然语言任务描述,即可自主生成多样化机器人运动原语档案的方法,在4项操作任务上优于经典QD算法。
AI 中文摘要
质量多样性(QD)算法在机器人学习领域正受到越来越多的关注,多样化的运动原语库使机器人能够在部署时零样本适应各种约束条件。然而,这类方法通常需要专家设计师编写成功条件、适应度和多样性指标,这极大地限制了机器人的自主性。另一方面,现有的基于大语言模型(LLM)的奖励塑形技术允许机器人自主学习,但仅能输出单一高性能解决方案,限制了机器人的适应性。在本文中,我们提出一种方法,旨在通过自主利用质量多样性算法输出多样化的运动原语档案,仅需以自然语言提供任务的自由形式描述。为解决设计相关适应度和多样性指标的难题,我们提出一种自主探索机制,能够可靠输出覆盖适应度和行为描述符(BD)空间的函数集。首先,我们将策略探索建模为函数设计问题,其中函数空间的维度低于完整的BD和适应度空间,并提出一种基于LLM的探索方案,无需任何任务特定提示、微调或专家干预即可从这些低维空间采样。我们适配了MAP-Elites成功(MES)算法的多BD变体,该算法旨在利用异构BD样本。最后,通过基于Genesis模拟器的实验,我们表明我们的方法能有效生成多样化运动原语档案,在4项机器人操作任务上,其性能优于具有推断和手动编写参数化的经典QD算法。
英文摘要
Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive libraries allow robots to adapt zero-shot to constraints at deployment time. However, such methods typically require expert designers to write the success condition, fitness and diversity metrics, and this strongly limits the robot's autonomy. On the other hand, existing LLM-based reward-shaping techniques allow robots to learn autonomously but only output single high-performing solutions, limiting the robot's adaptability. In this paper, we propose an approach designed to output diverse motion primitive archives by autonomously leveraging quality-diversity algorithms, only requiring a free-form description of the task in common language. To address the difficulty of designing relevant fitness and diversity metrics, we propose an autonomous exploration mechanism able to reliably output sets of functionals covering the fitness and behavior descriptor (BD) space. First, we pose policy exploration as a functional design problem, where the functional spaces are lower-dimensional than the full BD and fitness spaces, and propose an LLM-based exploration scheme to sample from these low-dimensional spaces without any task-specific prompts, fine-tuning or expert intervention. We adapt a multi-BD variant of the MAP-Elites success (MES) algorithm, designed to leverage the heterogeneous BD samples. Finally, through experiments based on the genesis simulator, we show that our method effectively generates archives of diverse motion primitives, outperforming classical QD algorithms with inferred and hand-written parametrizations on a set of $4$ robotic manipulation tasks.