AI 中文总结
本文提出一种基于强化学习的机器人建造方法HSAC,无需预定义计划,通过图神经网络和混合动作空间自适应生成建造序列,在仿真和物理双机器人实验中成功构建跨越拱门,性能优于HPPO。
AI 中文摘要
机器人建造有望更高效地利用材料并创造复杂几何形状,但当前方法依赖于刚性、高精度的计划,无法适应物理制造中固有的公差、误差和意外变化。在这项工作中,我们提出了一种强化学习方法,完全摒弃预定义计划,而是在结构建造过程中自适应地生成建造序列。我们的方法基于图结构状态表示和混合(参数化)动作空间,需要同时进行离散的块选择与连续的放置参数决策。由于结构的稳定性模拟计算量巨大,我们开发了一种高效的探索策略,将单边边引入图神经网络,并将软演员-评论家(SAC)算法扩展到这种混合设置中。我们将我们的算法HSAC与先前的方法hybrid-PPO(HPPO)进行了评估比较,展示了显著更高的渐近性能和良好的样本效率。我们还证明了HSAC对超参数选择的鲁棒性及其探索能力,能够在性能不下降的情况下处理多达10个离散动作。最后,我们在一个物理双机器人设置上验证了我们的方法,在闭环执行中成功使用3D打印块构建了一座跨越拱门,证实了在仿真中训练的策略能够迁移到真实硬件上。
英文摘要
Robotic construction offers the potential to use materials more efficiently and create complex geometries, but current methods rely on rigid, high-precision plans that cannot accommodate the tolerances, inaccuracies, and unexpected changes inherent in physical fabrication. In this work, we introduce a reinforcement learning approach that forgoes predefined plans entirely, instead generating construction sequences adaptively as the structure is built. Our method operates on graph-structured state representations and a mixed (parameterized) action space, requiring both discrete block selection and continuous placement parameters. Because the stability simulation of a structure is computationally heavy, we develop an efficient exploration strategy by incorporating unilateral edges into graph neural networks, extending soft actor-critic (SAC) to this hybrid setting. We evaluate our algorithm, HSAC, against the prior method hybrid-PPO (HPPO), demonstrating significantly higher asymptotic performance and good sample efficiency. We also demonstrate HSAC's robustness to hyperparameter choices and its exploration capability, handling up to 10 discrete actions without performance degradation. Finally, we validate our approach on a physical two-robot setup, successfully building a spanning arch with 3D-printed blocks in closed-loop execution, confirming that policies trained in simulation transfer to real hardware.
CommentsAccepted in IROS 2026