在仿真可变形表面上的演示引导类人机器人站立
Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface
浏览论文内容
中文总结 AI 辅助
该研究提出参考引导强化学习框架,利用硬地面人类演示,使29自由度Unitree G1类人机器人在仿真可变形软地面成功完成站立任务,明确恢复奖励是成功的关键。
中文摘要 AI 辅助
本文提出一种参考引导强化学习框架,用于在可变形软地面上为29自由度的Unitree G1类人机器人生成站立动作,该框架利用在硬地面上记录的人类演示。地形顺应性通过MuJoCo的刚体软接触模型中的solref和solimp参数进行建模。奖励由两部分组成:(i)通过残差关节位置控制实现的参考动作跟踪;(ii)明确的恢复目标,如骨盆高度、躯干直立度和最终姿势。首先,在考虑硬地面的情况下,使用指定奖励训练策略。接着,通过更新solref降低地形刚度,并使用solimp扩大名义表面穿透区域。后续训练使策略能够适应接触密集阶段因显著表面穿透而产生的延迟支撑力生成,同时保留原始演示模式。学习到的策略在仿真中成功完成从跌倒到站立的任务,达到目标骨盆高度和直立度,过程中最大接触穿透约为40毫米。所提出的方法在两个站立序列上进行了验证,在硬地面和软地面上均成功实现最终恢复目标。消融研究表明,仅参考跟踪不足以成功站立,明确的恢复奖励是必不可少的。
英文摘要
This paper presents a reference-guided reinforcement learning framework to generate stand-up motion for a 29-DOF Unitree G1 humanoid on deformable soft ground, using a human demonstration recorded on hard ground. The terrain compliance is modelled using solref and solimp parameters from MuJoCo's rigid body soft-contact model. The rewards consists of (i) reference motion tracking through residual joint-position control and (ii) explicit recovery objectives such as pelvis height, torso uprightness, and the final posture. First, the policy is trained with the specified rewards considering hard ground. Next, the terrain stiffness is lowered by updating solref and the nominal surface penetration zone is expanded using solimp. Subsequent training enables the policy to adapt to the delayed support force generation due to significant surface penetration during contact-intensive phases while preserving the original demonstration pattern. The learned policy successfully completes the fallen-to-standing task in simulation, reaching the targeted pelvis height and uprightness, with a maximum contact penetration of approximately 40 mm during the process. The proposed method is demonstrated on two stand-up sequences and successfully achieves the final recovery objective on both hard and soft ground. Ablation studies show that reference tracking alone is insufficient for successful stand-up, and that explicit recovery rewards are essential.
发表机构
- Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)
机构由 AI 辅助整理,请以论文原文为准。