发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分布式灵巧操作,提出基于空间条件多智能体Transformer的框架,结合自适应层归一化、空间对比嵌入和SAC微调,在64个软体Delta机器人阵列上实现长时域操作,动作选择减少65%机器人且误差约1.5厘米。
AI 中文摘要
分布式灵巧操作(DDM)是一种新颖的范式,由于高动作空间冗余、机器人间协作以及动态的物体-机器人交互,带来了显著的控制挑战。本文提出了一种基于空间条件多智能体Transformer(MATs)的框架,以高效学习针对DDM系统的鲁棒控制策略,该系统由排列在8x8网格中的64个软体Delta机器人组成。我们的三个核心贡献是:(i)具有自适应层归一化的MAT,以提高计算效率;(ii)空间对比嵌入,将Transformer嵌入锚定在机器人的空间配置中;(iii)基于MAT的行为克隆方法,并使用软演员-评论家(Soft Actor Critic)进行微调。我们还提出了一种动作选择公式,以分析任务性能与所用机器人数量之间的权衡。我们的实验表明,MAT通过堆叠的注意力块迭代地细化其动作。这进一步说明了Transformer中空间条件对学习DDM策略的益处。我们在仿真和真实世界中演示了具有各种几何形状物体的长时域平面操作任务。最后,我们展示了动作选择如何通过减少因机器人间碰撞造成的磨损来缓解机器人维护问题,同时保持沿各种轨迹操作物体的能力,在真实世界中实现了约1.5厘米的平均误差,同时使用了约65%更少的机器人。
英文摘要
Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due to high action-space redundancy, inter-robot cooperation, and dynamic object-robot interactions. This paper introduces a framework based on spatially conditioned Multi-Agent Transformers (MATs) to efficiently learn robust control policies for a DDM system grounded in an array of 64 soft delta robots arranged in an 8x8 grid. Our three core contributions are: (i) an MAT with adaptive layer norm for compute efficiency, (ii) spatial contrastive embeddings to ground transformer embeddings in the spatial configuration of the robots, and (iii) an MAT-based behavior cloning method fine-tuned using Soft Actor Critic. We also propose an action selection formulation to analyze the trade-off between task performance and the number of robots utilized. Our experiments show that MATs iteratively refine their actions through the stacked attention blocks. This further informs the benefit of spatial conditioning in transformers to learn DDM policies. We demonstrate long-horizon planar manipulation tasks with objects of various geometries in simulation and real-world. Finally, we show how action selection mitigates robot maintenance by reducing wear and tear due to inter-robot collisions while maintaining the ability to manipulate objects along various trajectories in the real-world, achieving an average error of ~1.5 cm, while using ~65% fewer robots.