arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22075cs.RO

LIMBO:学习并内化无模型障碍函数目标以实现敏捷且安全的全身控制

LIMBO: Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control

Jake Gonzales, Arturo Flores Alvarez, Yu-Ming Chen, Aaron D. Ames, Lillian J. Ratliff, Manikantan Nambi

首次发表
浏览论文内容

中文总结 AI 辅助

LIMBO框架通过风险引导采样学习Q-CBF安全证书并蒸馏至任务策略,在29自由度人形机器人上实现无需在线过滤的敏捷安全全身控制。

中文摘要 AI 辅助

安全的全身控制需要在高维、非线性动力学下协调避碰与平衡,这使得安全证书难以设计且难以在不同行为间复用。我们提出LIMBO,一个用于合成状态-动作控制障碍函数并将其安全结构蒸馏到任务策略中的框架。LIMBO从黑盒转移和基于状态的失败规范中学习安全证书,该规范针对冻结基础控制器周围的残差动作,从而在完整控制维度上使Q-CBF合成易于处理,同时将证书置于任务策略的控制空间中。在合成过程中,学习到的安全值驱动风险引导采样,使其接近可恢复性的估计边界;在任务学习过程中,它充当教师,提供动作级安全反馈,产生稳健的任务策略,并减轻部署时对在线安全过滤器的需求。我们在一个29自由度的人形机器人上演示了LIMBO,执行躲避球避让和在低障碍物下行走。除了将学习的Q-CBF扩展到全身控制外,我们表明风险引导的边界采样提供了一种理论上合理的方式来探索可恢复性的边缘。在相同安全规范下,保持其他条件不变,改变采样浓度会产生从蹲伏到新颖的后仰式limbo动作等不同策略。在两种设置中,学习到的策略无需在线安全过滤即可迁移到硬件,表明学习的安全合成可扩展到敏捷的全身控制。

英文摘要

Safe whole-body control requires coordinating collision avoidance and balance under high-dimensional, nonlinear dynamics--making safety certificates difficult to design and reuse across behaviors. We present LIMBO, a framework for synthesizing a state-action control barrier function and distilling its safety structure into a task policy. LIMBO learns the safety certificate from black-box transitions and a state-based failure specification over residual actions around a frozen base controller, making Q-CBF synthesis tractable in the full control dimension while placing the certificate in the task policy's control space. During synthesis, the learned safety value drives risk-guided sampling near the estimated boundary of recoverability; during task learning, it serves as a teacher that provides action-level safety feedback, yielding a robust task policy and alleviating the need for an online safety filter at deployment. We demonstrate LIMBO on a 29-degree-of-freedom humanoid performing dodgeball avoidance and locomotion beneath low obstacles. Beyond scaling learned Q-CBFs to whole-body control, we show that risk-guided boundary sampling provides a theoretically grounded way to explore the edge of recoverability. Under the same safety specification, ceteris paribus, varying the sampling concentration produces strategies ranging from crouching to a novel backward-leaning limbo maneuver. In both settings, the learned policies transfer to hardware without online safety filtering, showing that learned safety synthesis scales to agile whole-body control.

发表机构

  • Amazon(亚马逊)
  • University of Washington(华盛顿大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • California Institute of Technology(加利福尼亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑