arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

帕累托最优熵正则化轨迹优化

Pareto-Optimal Entropy-Regularized Trajectory Optimization

Dimitrios S. Georgiou, Augustinos D. Saravanos, Evangelos A. Theodorou

arXiv 2610.08532首次发表:更新:

发表机构

Georgia Institute of Technology; Massachusetts Institute of Technology(佐治亚理工学院; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非线性轨迹优化易陷入局部最优的问题,提出帕累托最优熵正则化DDP(PER-DDP),结合先验引导采样、扩展评估与帕累托过滤,在多个系统中超越现有方法,成功率更高且能解决基线无法处理的环境。

AI 中文摘要

在非线性动力学、执行器限制和避碰约束下的轨迹优化(TO)是机器人学中的一个基本问题,尽管由于其高度非凸的特性而特别具有挑战性。针对这一设定,微分动态规划(DDP)是一种高效的二阶打靶法,但其局部结构使其容易陷入次优盆地。采样增强的变体通过随机探索来缓解这种脆弱性,但通常仅围绕它们为重新优化而保留的少数轨迹进行采样,且仅基于其成本,这限制了探索广度。我们引入了帕累托最优熵正则化DDP(PER-DDP),这是一个源自自由能/相对熵不等式的熵正则化种群框架。我们的方法结合了先验引导的采样(它围绕每条保留的轨迹塑造探索)、扩展的 rollout 评估(它探测这些采样策略,超出少数保留的候选),以及帕累托过滤(用于在迭代中保留任务约束的替代方案)。这使采样工作与优化种群大小解耦,并在不牺牲使DDP有效的二阶结构的情况下拓宽了探索。在多个系统和数百个环境中,PER-DDP比最先进的采样增强TO方法实现了更高的成功率,并在所有基线无法企及的环境中找到了可靠的解决方案。

英文摘要

Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature. For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins. Sampling-augmented variants mitigate this susceptibility through stochastic exploration, but often sample only around the few trajectories they retain for reoptimization, based solely on their cost which restricts exploration breadth. We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality. Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations. This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective. Across multiple systems and hundreds of environments, PER-DDP achieves higher success rates than state-of-the-art sampling-augmented TO methods and finds reliable solutions in environments beyond the reach of all baselines.

Comments9 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑