arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

捆绑接触梯度:稳定可微仿真以支持可部署的动态任务

Bundled Contact Gradients: Stabilizing Differentiable Simulation for Deployable Dynamic Tasks

Dyuman Aditya, Jin Cheng, Clemens Schwarke, Quan Nguyen, Gaurav Sukhatme, Stelian Coros, Gabriele Fadini

arXiv 2609.30951首次发表:更新:

发表机构

University of Southern California; ETH Zurich; NVIDIA; Zurich University of Applied Sciences(南加州大学; 苏黎世联邦理工学院; 英伟达; 苏黎世应用科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对刚性接触导致的可微仿真梯度高方差问题,提出捆绑接触梯度(BCG)框架,通过局部随机平滑降低梯度方差,成功训练并零样本迁移动态动作至真实Unitree G1人形机器人。

AI 中文摘要

可微仿真提供了机器人动力学的解析梯度,使得快速且样本高效的一阶策略优化成为可能。然而,通过刚体接触获得平滑且信息丰富的梯度通常需要软化接触模型,这往往以物理保真度为代价,从而将学习到的策略在很大程度上限制在仿真环境中。这种权衡对于动态人形运动尤为重要,因为准确的接触动力学对于将策略迁移到现实世界至关重要。在刚体仿真中增加接触刚度可以提高交互的保真度,但也会使动力学对小状态扰动越来越敏感,产生高方差梯度,从而可能破坏一阶策略学习的稳定性。为了解决这一问题,我们提出了捆绑接触梯度(BCG),一种用于可微策略学习的接触局部随机平滑框架。当检测到刚性接触时,我们的方法在刚性接触配置周围评估一组局部随机扰动回滚,并聚合其梯度信号,从而降低梯度方差。我们通过成功训练并将动态动作零样本迁移到真实的Unitree G1人形平台上,证明了我们方法的有效性。视频和补充信息可在该https URL找到。

英文摘要

Differentiable simulation provides analytic gradients of robot dynamics, enabling fast and sample-efficient first-order policy optimization. However, obtaining smooth and informative gradients through rigid-body contact typically requires softened contact models, often at the expense of physical fidelity and thereby limiting learned policies largely to simulation. This trade-off becomes particularly consequential for dynamic humanoid motions, where accurate contact dynamics are critical for transferring policies to the real world. Increasing contact stiffness in rigid-body simulation improves the fidelity of interactions, but also makes the dynamics increasingly sensitive to small state perturbations, producing high-variance gradients that can destabilize first-order policy learning. To address this, we propose \emph{Bundled Contact Gradients (BCG)}, a contact-local randomized smoothing framework for differentiable policy learning. When stiff contact is detected, our method evaluates a local bundle of randomized perturbation rollouts around the stiff contact configuration and aggregates their gradient signal thereby reducing gradient variance. We demonstrate the effectiveness of our method by successfully training and transferring dynamic motions zero-shot onto a real-world Unitree G1 humanoid platform. Videos and supplementary information can be found at https://bundledcontactgradients.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑