发表机构
Technical University of Darmstadt; Robotics Institute Germany (RIG); LimX Dynamics; German Research Center for AI (DFKI)(达姆施塔特工业大学; 德国机器人研究所; LimX Dynamics; 德国人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VAPS通过策略条件化滚动时域决策,在跟踪、中止和摔倒策略间动态选择,显著减少人形机器人翻跟头时的硬件损坏风险。
AI 中文摘要
动态人形运动(如翻跟头)因次优策略、扰动或仿真到现实的差距而面临硬件损坏风险。一旦动作偏离参考轨迹,运动跟踪策略便无计可施,此时需要备用策略接管以最小化损坏落地。选择哪个备用策略与何时切换同样重要。我们提出生存能力感知策略选择(VAPS),将安全视为策略条件化的滚动时域决策。除保护性摔倒策略外,我们还训练了一个中止策略,可随时中止动作并双脚落地。在每个控制步,学习型预测器估计名义跟踪策略和中止策略在短时域内是否仍具生存能力,一个最小牺牲层级结构保留最具任务雄心且仍具生存能力的行为。在带随机扰动的仿真中,VAPS显著减少了头部和手部接触——这是硬件损坏的主要来源,在Unitree G1和LimX Oli上均得到验证;在LimX Oli上,我们验证了侧翻动作的生存能力预测器和完整VAPS控制器。VAPS在任务成功率和头部冲击方面帕累托支配了我们能训练的最强单网络替代方案,包括端到端安全跟踪策略和从VAPS自身预言机路由决策中蒸馏的学生策略。我们还表明,VAPS是监督训练不足策略并保护硬件的强大框架。
英文摘要
Dynamic humanoid motions such as flips risk hardware damage due to suboptimal policies, disturbances or sim-to-real gaps. A motion tracking policy offers no way out once the maneuver leaves its reference, and a backup policy needs to take over to protect the hardware for a minimum-damage landing. Which backup to use matters as much as when to switch. We present Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision. Besides a protective fall policy, we also train an abort policy which can abort the motion at any time, landing on its feet. At every control step, learned predictors estimate whether the nominal tracking policy and the abort policy remain viable over a short horizon, and a least-sacrificial hierarchy keeps the most task-ambitious behavior that remains viable. In simulation with randomized disturbances, VAPS sharply reduces head contact and hand contact, which are the dominant sources of hardware damage, with both a Unitree G1 and a LimX Oli; on the LimX Oli, we validate the viability predictors and the full VAPS controller for side-flip motions. VAPS Pareto-dominates the strongest single-network alternatives we could train, including an end-to-end safe-tracking policy and students distilled from VAPS's own oracle-routed decisions, in both task success and head impact. We also show that VAPS is a powerful framework to supervise undertrained policies and protect the hardware.