arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Box-iLQR:通过有限障碍轨迹优化实现约束感知反馈

Box-iLQR: Constraint-Aware Feedback via Finite-Barrier Trajectory Optimization

Abhijeet, Suman Chakravorty

arXiv 2610.04712首次发表:更新:

发表机构

Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Box-iLQR,一种针对盒式约束的对数障碍iLQR方法,通过保留有限障碍曲率来塑造反馈增益,在约束边界处保留控制裕度,实验表明其闭环性能优于基线方法。

AI 中文摘要

迭代线性二次调节器(iLQR)和微分动态规划(DDP)是极具吸引力的轨迹优化方法,因为它们的反向递推既能提供名义轨迹,也能提供局部时变反馈律。然而,硬状态和输入约束可能将解推向约束边界,在闭环执行过程中几乎没有剩余的控制修正能力。本文提出了Box-iLQR,一种针对盒式约束状态和控制的对数障碍iLQR方法,重点研究由此产生的反馈策略。障碍梯度使名义轨迹偏离约束边界,而障碍曲率进入决定反馈增益的Riccati递推中。我们证明,控制障碍曲率在执行器接近饱和时平滑地衰减反馈,而状态障碍曲率通过值函数向后传播,在未来状态约束遇到之前塑造反馈。这促使保留有限障碍而不是总是将障碍参数推向零,从而保留状态和执行器裕度。在摆、小车-倒立摆、Acrobot和汽车停车问题上,与CL-DDP和ALTRO进行了比较。蒙特卡洛实验证明了有限障碍下改进的闭环行为,而一个六输入MuJoCo鱼示例则展示了使用从模拟器滚动估计的动力学雅可比矩阵的数据驱动Box-iLQR。Box-iLQR实现了与基线方法相当甚至在某些情况下更好的收敛性。此外,使用有限障碍获得的反馈策略即使在执行不确定性下也能确保约束满足。

英文摘要

Iterative Linear Quadratic Regulator (iLQR) and Differential Dynamic Programming (DDP) are attractive trajectory-optimization methods because their backward recursions provide both nominal trajectories and local time-varying feedback laws. Hard state and input constraints, however, can drive solutions toward constraint boundaries where little corrective authority remains during closed-loop execution. This paper presents Box-iLQR, a log-barrier iLQR method for box-constrained states and controls, with emphasis on the resulting feedback policy. Barrier gradients bias nominal trajectories away from constraint boundaries, while barrier curvature enters the Riccati recursion that determines the feedback gains. We show that control-barrier curvature smoothly attenuates feedback as an actuator approaches saturation, while state-barrier curvature propagates backward through the value function to shape feedback before future state constraints are encountered. This motivates retaining a finite barrier rather than always driving the barrier parameter toward zero, thereby preserving state and actuator margin. Comparisons with CL-DDP and ALTRO are presented on pendulum, cart-pole, Acrobot, and car-parking problems. Monte Carlo experiments demonstrate improved closed-loop behavior with finite barriers, while a six-input MuJoCo fish example demonstrates data-driven Box-iLQR using dynamics Jacobians estimated from simulator rollouts. Box-iLQR achieves convergence comparable to, and in some cases better than, the baseline methods. Moreover, the feedback policy obtained with a finite barrier ensures constraint satisfaction even under execution uncertainty.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑