LTLDiff:有限线性时序逻辑引导的数据生成与扩散策略用于多智能体机器人操作
LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation
- University of Toronto(多伦多大学)
- A*STAR Centre for Frontier AI Research (A*STAR CFAR), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局前沿人工智能研究中心(A*STAR CFAR))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多智能体操作中的协调失败问题,提出LTLDiff框架,结合有限线性时序逻辑规范学习与扩散策略,实现逻辑引导的数据生成和策略训练,提升任务成功率。
AI中文摘要:
多智能体机器人操作任务要求智能体之间进行协调,以满足任务层面的时序、逻辑和安全约束。近年来,扩散策略已被用于执行此类任务。然而,在需要同时或顺序多智能体交互的任务中,这些策略仍存在失同步、动作顺序错误和协调失败等问题。为此,本文提出LTLDiff框架,该框架结合了有限线性时序逻辑(LTLf)规范学习,用于演示生成和基于扩散策略的学习。每个任务都有一个特定的LTLf公式,该公式通过大规模语言模型从一组自然语言指令中学习得到。为了实现对语言模型学习到的规范进行固定维度向量嵌入,LTLf采用抽象语法树表示方案。这种逻辑嵌入作为(i)逻辑引导的数据收集和(ii)基于扩散的策略训练的条件,鼓励生成符合期望顺序和协调要求的轨迹。在多智能体LTLDiff操作任务上的实验表明,与基线相比,任务成功率有所提高。综合来看,这些贡献证明了LTLDiff在协调多智能体操作方面的有效性。
英文摘要:
Multi-agent robotic manipulation tasks require coordination among agents to satisfy task-level temporal, logical, and safety constraints. Recently, diffusion policies have been used to perform the task. However, they still suffer from desynchronization, incorrect action ordering, and coordination failures in tasks that require simultaneous or sequential multi-agent interaction. Therefore, LTLDiff is proposed as a framework that combines Finite Linear Temporal Logic (LTLf) specification learning for both the generation of demonstrations and learning via diffusion policies. Each task has a specific LTLf formula that is learned from a set of natural language instructions using a large-scale language model. To enable a fixed-dimensional vector embedding of the learned specification from the language model, LTLf uses an abstract syntax tree representation scheme. This embedding of logic serves as a condition for (i) logic-guided data collection and (ii) diffusion-based policy training, encouraging trajectories that are consistent with the desired ordering and coordination requirements. Experiments on multi-agent LTLDiff manipulation tasks demonstrate improved task success rates compared to the baseline. Together, these contributions demonstrate the effectiveness of LTLDiff for coordinated multi-agent manipulation.