通过多目标强化学习实现四足机器人运输无固定堆叠载荷
Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning
- Fujitsu Limited(富士通株式会社)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对四足机器人运输无固定堆叠载荷时运动与稳定性的权衡,提出PAMORT多目标强化学习方法,通过在线调整偏好权重,在仿真和真实实验中优于单目标基线,实现鲁棒运输。
AI中文摘要:
使用腿式机器人在不平坦地形上运输无固定载荷需要在运动性能与载荷稳定性之间取得平衡,因为即使机器人保持稳定,剧烈的运动也可能使载荷失稳。我们研究了在无专用载荷传感器或主动承载机构的情况下,四足机器人运输放置在无边缘躯干安装板上的无固定堆叠箱体的问题。为解决这一权衡问题,我们提出了载荷自适应多目标强化学习运输方法(PAMORT)。PAMORT 训练一个以偏好向量为条件的多目标基础策略,该偏好向量对运动与载荷稳定性奖励组进行加权,然后在冻结策略上训练一个权重调整器,以根据本体感觉在线调整该偏好。在仿真中,与相应的单目标基线相比,PAMORT 在不同载荷配置(包括未见过的三箱堆叠)下实现了相当或更好的整体运输成功率,尽管仅使用两箱进行训练。在 Unitree Go2 上的真实世界实验展示了在达到或超过训练难度的斜坡和台阶上的零样本迁移,在八项任务中 PAMORT 的平均成功率为 0.850,基线为 0.675。这些结果表明,通过本体感觉信息在线调整运动-载荷权衡,实现了鲁棒的无固定载荷运输。
英文摘要:
Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.