arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习双足移动操作机器人的全身整体移动操作

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

Zhongyu Chen, Yuxuan Nai, Qian Chen, Yidong Zhu, Chen Jing, Qihan Wang, Xudong Li, Zhizhan Li, Leixin Chang, Liangjing Yang, Hua Chen

arXiv 2609.18930首次发表:更新:

发表机构

ZJU-UIUC Institute; LimX Dynamics(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合学院; LimX Dynamics(深圳逐际动力))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于强化学习的统一全身控制器,直接映射6自由度末端执行器目标,协调双足移动操作机器人的伸展、姿态适应与迈步,并通过真实机器人实验验证其通用性。

AI 中文摘要

双足移动操作使机器人能够通过协调移动和操作,与手臂标称工作空间之外的物体进行交互。实现这一能力需要一个低层全身控制器,将任务级操作目标转化为协调的手臂和腿部运动,同时保持平衡。我们提出了一种使用强化学习训练的统一全身控制器,该控制器直接将6自由度末端执行器目标映射为双足基座和机械臂的协调动作。仅给定末端执行器目标,学习到的控制器自主协调伸展、姿态适应和迈步,无需显式的基座速度或脚步命令。在训练过程中,一种奖励门控策略调节末端执行器跟踪、移动和平衡之间的权衡,而时间上下文估计器结合了窗口化Transformer编码、循环GRU记忆和辅助动力学预测,从观测历史中提取动力学相关信息。真实机器人实验表明,同一控制器支持在VR遥操作、学习到的扩散策略和脚本化轨迹命令下的伸展、姿态适应和迈步,为多样化操作任务提供了通用的末端执行器接口。

英文摘要

Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑