arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24959cs.ROcs.DCcs.LGcs.SYeess.SYmath.OC

通过隐式接触微分实现用于残差MPC的摊销轨迹优化

Amortising Trajectory Optimisation for Residual MPC via Implicit Contact Differentiation

Daniel Layeghi, Thomas Corbères, Calum Arnott, Aditya Kamireddypalli, Hashim Al-Obaidi, Steve Tonneau, Michael Mistry

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对残差MPC的轨迹优化问题,基于隐函数定理引入AD辅助隐式导数用于正则化平滑接触,避免求解器展开和复杂推导,还引入优化器蒸馏策略,相比标准iLQR显著提高了相关任务的成功率。

中文摘要 AI 辅助

可微模拟通过揭示任务结果对控制的局部敏感性来加速富含接触的轨迹优化。现有方法要么使用有限差分,成本高且对步长敏感;要么通过展开自动微分来区分迭代接触求解器,这会存储不断增长的计算轨迹;要么需要复杂的、特定求解器的KKT灵敏度推导。我们基于隐函数定理为正则化平滑接触引入了一种AD辅助隐式导数,并将其应用于Mujoco MJX。该方法在容差收敛解处区分平稳残差,避免了求解器展开和手工组装的KKT系统。隐函数定理使编译后的临时内存与求解器工作量几乎保持不变,从一次迭代到十次迭代变化小于4%,而展开的自动微分增长10.6倍。随着活动接触和模型维度的增加,隐函数定理的内存增长更慢,在256个接触点时使用的内存减少20倍,在16个接触点和96自由度时减少6倍。我们还为残差MPC引入了优化器蒸馏,将批量全时域iLQR摊销到一个指导短时域残差iLQR的策略中。在Finger、Franka和Unitree上测试,这比标准iLQR将六步成功率提高了28 - 98个百分点。

英文摘要

Differentiable simulation can accelerate contact-rich trajectory optimisation by exposing local sensitivities of task outcomes to controls. Existing approaches either use finite differences, which are expensive and step-size sensitive; differentiate iterative contact solvers by unrolling automatic differentiation (AD), which stores a growing computation trace; or require intricate, solver-specific KKT sensitivity derivations. We introduce an AD-assisted implicit derivative for regularised smooth contacts and apply it to Mujoco MJX, based on the Implicit Function Theorem (IFT). The method differentiates the stationarity residual at the tolerance-converged solution, avoiding both solver unrolling and hand-assembled KKT systems. IFT keeps compiled temporary memory nearly constant with solver effort, changing by less than 4$\%$ from one to ten iterations versus 10.6$\times$ growth for unrolled AD. IFT memory grows slower with active contacts and model dimension, using 20$\times$ less memory at 256 contacts and 6$\times$ less at 16 contacts and 96 DoF. We further introduce optimiser distillation for residual MPC, amortising batched full-horizon iLQR into a policy that guides short-horizon residual iLQR. Across Finger, Franka, and Unitree, this raises six-step success by 28-98 percentage points over standard iLQR.

发表机构

  • School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑