arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20823cs.RO

通过引导式辅助课程与分阶段奖励实现类人机器人的自然坐站运动合成

Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards

Meet Pal Singh, Vyankatesh Ashtekar, Ashish Dutta

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对类人机器人坐站控制难题,提出结合耦合力/椅子高度课程、鲁棒采样及生物力学启发奖励函数的强化学习方法,实现多椅子高度下97%以上平衡站立成功率,泛化性优于现有技术。

中文摘要 AI 辅助

类人机器人从坐姿站立时,在保持平衡的前提下存在无限多种方式,这使得坐站(STS)成为一项极具挑战性的控制问题。本文使用强化学习从零开始合成类人机器人的自然坐站运动,无需演示或参考轨迹。单个近端策略优化(Proximal Policy Optimisation)策略通过三个互补组件学习流畅、类人的起身动作:(i)采用耦合力/椅子高度课程,垂直骨盆辅助力助力早期轨迹探索并在训练过程中衰减,随着辅助力衰减解锁更高的椅子,确保策略在每个椅子高度掌握可行的坐站轨迹后再接触更难的高度,避免过早分布偏移导致泛化崩溃;(ii)通过从大量逆运动学生成的初始和目标姿态中随机采样实现运动鲁棒性,这些姿态覆盖超过8种椅子高度;(iii)受生物力学和最优控制研究启发定义了一组奖励函数,它们塑造机器人离地时的角动量,并通过压力中心吸引函数实现支撑区域过渡,确保低能耗的流畅驱动。在确定性无力量化评估器上,该策略在8种椅子高度下均实现超过97%的平衡站立成功率,相比现有技术,该策略可在不同椅子高度间泛化流畅运动,使机器人能从显著更深的坐姿起身。

英文摘要

A humanoid has infinitely many ways to stand up from sitting while maintaining balance, making sit-to-stand (STS) a challenging control problem. We synthesise natural humanoid STS motion from scratch using reinforcement learning, without demonstrations or reference trajectories. A single Proximal Policy Optimisation policy learns smooth, human-like rising driven by three complementary components. (i) A coupled force/chair-height curriculum is used. A vertical pelvis-assist force aids early trajectory exploration and decays over training. Taller chairs are unlocked with decaying assisting force. This ensures that the policy masters a viable STS trajectory at each chair height before being exposed to harder ones, avoiding the premature distribution shift that otherwise collapses generalisation. (ii) Motion robustness is achieved by randomly sampling from a large number of inverse kinematics-generated initial and target poses spanning over eight chair heights. (iii) A set of rewards is defined inspired from biomechanics and optimal control studies. They shape the robot's angular momentum for seat-off, and enable support-region transition via centre of pressure attraction function to ensure smooth low-effort actuation. On a deterministic force-free evaluator, the policy attains more than 97% balanced-standing success across eight chair heights. The policy generalises smooth motion across chair heights and enables the robot to rise from substantially deep-seated postures as compared to the state of the art.

发表机构

  • Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

↑