arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向安全人形全身跟踪的滤波器感知微调

Filter-Aware Fine-Tuning for Safe Humanoid Whole-Body Tracking

Pranit Mohnot, Christian Helten, Daniele Gammelli, Marco Pavone

arXiv 2610.02341首次发表:更新:

发表机构

Stanford University; Technical University of Munich; Italian Institute of Artificial Intelligence (AI4I); NVIDIA(斯坦福大学; 慕尼黑工业大学; 意大利人工智能研究所(AI4I); 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对人形机器人全身跟踪,提出滤波器感知微调方法CoFiT,解决策略与安全滤波器不匹配问题,显著降低违规时间并提升硬件部署安全性。

AI 中文摘要

安全的全身体运动对于在非结构化环境中部署人形机器人至关重要。现代人形控制通常将参考规范与执行分离,由规划器、远程操作员或运动生成器提供参考,强化学习策略通过动态可行的全身控制来跟踪该参考。运行时安全滤波器,如控制屏障函数(CBFs),提供了一种有前景的方法,通过对跟踪器输出的干预来强制执行新引入的约束。然而,我们表明,将跟踪策略和安全滤波器独立处理会导致根本性的不匹配,因为滤波会改变执行的动作和诱导的状态分布。我们通过案例研究来研究这种策略-滤波器接口,这些案例研究隔离了动力学、目标和信息不匹配,突出了其根本原因,并利用这些见解开发了CoFiT(约束滤波器感知微调),一种针对预训练跟踪器的滤波器感知微调方法。在多样化的约束场景中,与仅滤波训练相比,CoFiT在TWIST2上将违规时间减少了91%,在SONIC上减少了21%,同时需要更小的安全滤波器修正。在Unitree G1硬件上,CoFiT将TWIST2的违规时间减少了83%,并且每次试验都在没有操作员干预的情况下完成,而50%的基线试验需要操作员停止。总之,这些结果为策略-滤波器交互提供了可操作的见解,并为将学习型跟踪器与运行时安全滤波器集成建立了设计原则。

英文摘要

Safe whole-body motion is essential for deploying humanoid robots in unstructured environments. Modern humanoid control commonly separates reference specification from execution, with a planner, teleoperator, or motion generator providing a reference that a reinforcement-learning policy tracks through dynamically feasible whole-body control. Runtime safety filters, such as control barrier functions (CBFs), offer a promising approach for enforcing newly introduced constraints via interventions on the tracker's outputs. We show, however, that treating the tracking policy and safety filter independently induces fundamental mismatches, as filtering alters both the executed actions and the induced state distribution. We study this policy-filter interface through case studies that isolate dynamics, objective, and information mismatches, highlight their root causes, and use these insights to develop CoFiT (Constrained Filter-aware Tuning), a filter-aware fine-tuning method for pretrained trackers. Across diverse constraint scenes, CoFiT reduces violation time relative to filter-only training by 91% on TWIST2 and 21% on SONIC, while requiring smaller safety filter corrections. On Unitree G1 hardware, CoFiT reduces violation time by 83% for TWIST2 and completes every trial without operator intervention, whereas 50% of baseline trials require an operator stop. Together, these results provide actionable insights into policy-filter interactions and establish design principles for integrating learned trackers with runtime safety filters.

Comments8 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑