arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人形机器人在受限空间中通过自碰撞规避实现全身规划的参考方案

Whole-Body Planning for Humanoids Navigating Confined Spaces via Self-Collision Avoidance References

Carlos Gonzalez, Luis Sentis

arXiv 2608.10220首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出三阶段全身规划框架,结合可微分自碰撞规避与可达性约束,在Unitree G1人形机器人上实现Cr<1.5的受限空间导航,生成复杂接触轨迹并训练出鲁棒在线跟踪策略。

AI 中文摘要

人形机器人在高度受限环境中移动时,需要在维持多接触动态可行性的同时,避开密集的环境障碍物与复杂的自碰撞边界。传统轨迹优化器在这类受限空间中常遇困难,因为在粒子抽象层面用样条曲线遍历大碰撞空间的方式存在不足,易陷入局部极小值。为此,本文提出一种三阶段全身规划框架,直接在运动学可达的刚体体积上构建运动学路径规划。通过将可微分自碰撞规避融入可达性约束的公式中,该框架生成由体积信息指导的参考轨迹,能可靠引导全阶轨迹优化器完成长时程规划。研究表明,这些优化后的规划可作为高质量参考,用于训练残差强化学习策略以实现鲁棒在线执行。我们在Unitree G1人形机器人上,针对三个符合NIST应急响应标准的基准测试平台验证了所提方法,实现了受限空间比率(Cr < 1.5)。在12至18秒的任务中,我们的框架能生成包含复杂足部与手部接触的可行轨迹,而标准基准方法均失效;同时,经训练的策略在物理仿真的广泛域随机化条件下可成功跟踪这些规划。

英文摘要

Humanoid locomotion in highly confined environments requires navigating dense environmental obstacles and complex self-collision bounds while maintaining multi-contact dynamic feasibility. Traditional trajectory optimizers frequently struggle in these restricted spaces, as navigating the large collision space with splines on particle abstractions is insufficient and leads to poor local minima. To address this, we propose a three-stage whole-body planning framework that formulates kinematic path planning directly over kinematically reachable rigid-body volumes. By integrating differentiable collision avoidance into a reachability-constrained formulation, our framework synthesizes volume-informed guides that reliably guide a full-order trajectory optimizer over long horizons. We show that these optimized plans serve as high-quality references to train a residual reinforcement learning policy for robust online execution. We validate our approach on the Unitree G1 humanoid across three benchmark testbeds exceeding NIST emergency response standards, achieving restricted confinement ratios ($C_r < 1.5$). Our framework generates feasible trajectories across 12-to-18-second tasks with complex foot and hand contacts where standard baselines fail, while the learned policy successfully tracks these plans under extensive domain randomization in physics simulation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑