arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08800cs.ROcs.GRcs.LG

Ostrich:在可微动力学中通过刚性接触大步前进

Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics

Aleš Kučera, Karel Zimmermann

首次发表
浏览论文内容

中文总结 AI 辅助

提出Ostrich,一种GPU加速的刚体模拟器,通过非光滑牛顿迭代和隐函数定理,在大时间步长下实现高精度可微接触,梯度可靠且内存高效,显著提升优化吞吐量。

中文摘要 AI 辅助

三个属性决定了一个可微模拟器能否通过接触驱动基于梯度的优化:模拟精度、梯度可靠性以及每次迭代的成本。基于磁带(tape-based)的引擎,如MJX和Newton Semi-Implicit,需要足够小的时间步长以保持接触在数值上可处理,并且它们的反向传播内存随时间步数T线性增长。代理模型通过近似接触来限制内存,但由此产生的梯度丢失了优化所依赖的几何信息。我们提出了Ostrich,一个GPU加速的刚体模拟器,它在大时间步长(h ~ 0.1秒)下通过非光滑牛顿迭代解决硬接触和摩擦问题,并通过隐函数定理对收敛残差进行微分,重用前向Schur补来计算伴随,每个时间步的内存为O(1)。在真实机器人穿越托盘障碍物的轨迹上,Ostrich在高达50倍更大的时间步长下保持了MuJoCo的仿真到现实精度。其梯度从随机初始化中收敛,而MJX下降缓慢,Newton Semi-Implicit停滞;一次预热迭代比MJX快211倍,比Semi-Implicit快4.7倍。在同一场景中,Ostrich在单个24 GB GPU上对8,192个并行世界进行微分,维持了29倍于检查点MJX的优化吞吐量;没有检查点的情况下,两个基线在远少于这些世界数时就耗尽了内存。最后,我们展示了在三角形网格地形上跨越10秒时域的基于梯度的轨迹优化演示,这是一个先前引擎要么局限于原始几何,要么面临上述收敛和内存限制的场景。

英文摘要

Three properties determine whether a differentiable simulator can drive gradient-based optimization through contact: simulation accuracy, gradient reliability, and per-iteration cost. Tape-based engines such as MJX and Newton Semi-Implicit require timesteps small enough to keep contacts numerically tractable, and their backpropagation memory grows linearly with the number of timesteps T. Surrogate models bound memory by approximating contact away, but the resulting gradients lose the geometry the optimization depends on. We present Ostrich, a GPU-accelerated rigid-body simulator that resolves hard contacts and friction with non-smooth Newton iteration at large timesteps (h ~ 0.1 s), and differentiates the converged residual via the implicit function theorem, reusing the forward Schur complement to compute the adjoint at O(1) memory per timestep. On real-robot trajectories over a pallet obstacle, Ostrich holds MuJoCo's sim-to-real accuracy up to a 50x larger timestep. Its gradients converge from random initializations where MJX descends slowly and Newton Semi-Implicit stalls; a warm iteration runs 211x faster than MJX's and 4.7x faster than Semi-Implicit's. On the same scene Ostrich differentiates 8,192 parallel worlds on a single 24 GB GPU, sustaining 29x checkpointed MJX's optimization throughput; without checkpointing both baselines exhaust memory at far fewer worlds. We close with a gradient-based trajectory optimization demonstration over triangle-mesh terrain across a 10 s horizon, a setting where prior engines either restrict to primitive geometry or face the convergence and memory limits shown above.

发表机构

  • Czech Technical University in Prague(布拉格捷克技术大学)
  • Faculty of Electrical Engineering(电气工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑