arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于序贯局部策略评估的随机多射击轨迹优化

Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation

Ashwin Gupta, Joseph Moore

arXiv 2608.03978首次发表:更新:

AI 中文总结

本文提出一种基于序贯局部策略评估的随机多射击轨迹优化方法,通过优化短控制动作序列提升样本效率与终端集收敛性,可用于黑盒动力学的基于模型强化学习,在三类欠驱动任务中验证了其有效性。

AI 中文摘要

随机单射击轨迹优化方法,如模型预测路径积分控制(MPPI),因能对概率动力学进行推理,且在模型梯度噪声大、评估成本高或不可用时仍能提供解决方案,已在机器人领域广泛应用。然而,对长动作序列进行射击时,终端约束的满足往往样本效率低下,需要大量迭代才能收敛。本文提出一种随机多射击方法,该方法优化通过局部反馈策略连接的短控制动作序列,以提高样本效率和终端集收敛性。此外,我们证明可仅通过 rollout(滚出)合成近似系统雅可比矩阵,使该方法适用于具有黑盒动力学的基于模型的强化学习。我们在三个非线性、欠驱动优化问题上验证了该算法,包括具有解析动力学的经典倒立摆摆起任务、具有学习到的神经网络动力学的倒立摆摆起任务,以及执行大攻角精确失速后着陆机动的垂直起降四轴飞行器,结果表明该算法样本效率和终端集收敛性均得到提升。

英文摘要

Stochastic single shooting trajectory optimization methods such as Model Predictive Path Integral control (MPPI) have been widely adopted in robotics due to their ability to reason about probabilistic dynamics and provide solutions where model gradients are noisy, costly to evaluate, or unavailable. However, satisfaction of terminal constraints when shooting over long action sequences is often sample inefficient, requiring a large number of iterations for convergence. In this paper, we present a stochastic multiple shooting method that optimizes short control action sequences connected via local feedback policies to improve sample efficiency and convergence to a terminal set. Additionally, we show that we are able to synthesize approximate system Jacobians purely from rollouts, making the method suitable for model-based reinforcement learning with black-box dynamics. We demonstrate the algorithm has improved sample efficiency and terminal set convergence for three nonlinear, underactuated optimization problems: a classic cartpole swingup task with analytical dynamics, a cartpole swingup task with learned neural network dynamics, and a VTOL quadplane performing a high angle-of-attack, precision post-stall landing maneuver.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑