arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04794math.OCcs.SYeess.SY

域随机化线性二次系统的策略迭代

Policy Iteration for Domain Randomized Linear Quadratic Systems

Abbas Pasdar, Farnaz Adib Yaghmaie

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对线性二次控制的域随机化问题,提出带稳定步长规则的策略迭代算法,证明其样本平均目标单调改进,在平滑假设下子收敛至驻点,梯度主导条件下全局线性收敛。

中文摘要 AI 辅助

在本研究中,我们针对线性二次控制的域随机化下的策略优化展开研究,重点学习单一状态反馈控制器,以最小化具有不确定动力学的系统的平均代价。我们提出一种带步长规则的策略迭代算法,该规则可在每次迭代中保持所有采样系统的稳定性。我们证明该方法能使样本平均目标函数单调改进,且始终存在稳定步长。在标准平滑假设下,迭代序列子收敛至驻点;在梯度主导条件下,我们得到全局线性收敛速率。

英文摘要

In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability across all sampled systems at each iteration. We show that the method yields monotonic improvement of the sample-average objective and that a stabilizing step size always exists. Under standard smoothness assumptions, the iterates converge subsequentially to stationary points, and under a gradient-dominance condition, we obtain a global linear convergence rate.

发表机构

  • Linköping University(林雪平大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑