发表机构
Örebro University(厄勒布鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人形机器人全身控制中平滑性与稳定性难以兼顾的问题,提出解耦约束感知策略DeCap,将平滑性作为上下半身显式物理约束,显著降低动作率与加速度,并跨地形迁移。
AI 中文摘要
具身人工智能系统,特别是在现实世界场景中部署的人形机器人,需要全身控制策略,这些策略既要对任务响应灵敏,又要在物理上保持平滑。然而,平滑性在身体各部位并不统一:下半身必须保持足够的反应性,而上半身必须受到严格调节以维持稳定性。现有的强化学习方法通常通过在奖励函数中加入辅助项来施加平滑性,这些辅助项与任务目标相竞争,将身体视为均匀一致,并且无法直接控制负责平滑行为的物理量。我们提出了DeCap(解耦约束感知策略),一种约束强化学习算法,它将全身平滑性解耦为独立的上半身和下半身约束组,每组将平滑性表述为对物理运动极限的显式约束。为了提高接近可行性边界时的约束满足度,DeCap引入了一种有界障碍惩罚,该惩罚在接近极限时主动激活,并在约束极限处保持有界。在真实世界的人形机器人全身控制任务中,与基于奖励的平滑性策略相比,DeCap将上半身动作率降低了2.50倍,加速度降低了2.18倍,同时改善了下半身平滑性并减少了瞬态运动。我们证明了固定的一组平滑性约束可以迁移到各种地形,从而减轻了广泛奖励调整的需求。
英文摘要
Embodied AI systems, particularly humanoid robots deployed in real world scenarios require whole-body control policies that are both task-responsive and physically smooth. However, smoothness is not uniform across the body: lower body must remain sufficiently reactive, while the upper body must be tightly regulated to preserve stability. Existing reinforcement learning approaches typically impose smoothness through auxiliary terms in the reward function, which compete with task objectives, treating the body as uniform and provide no direct control over the physical quantities responsible for smooth behavior. We introduce DeCap (Decoupled Constraint-aware policy), a constrained reinforcement learning algorithm that decouples whole-body smoothness into separate upper- and lower-body constraint groups, each formulates smoothness as explicit constraints on physical motion limits. To improve constraint satisfaction near feasibility boundaries, DeCap incorporates a bounded barrier penalty that activates proactively as limits are approached and remains bounded at the constraint limit. On real-world humanoid whole-body control task, DeCap reduces upper-body action rate by 2.50x and acceleration by 2.18x relative to reward-based smoothness policies, while also improving lower-body smoothness and reducing transient motion. We demonstrate that a fixed set of smoothness constraints transfers across diverse terrains, alleviating the need of extensive reward tuning.