AI 中文总结
该研究针对增益调度LQR策略优化,建立了代价关于极小值点的星形凸性恒等式,明确了收敛的无量纲比率阈值,通过实验验证了阈值为代价景观的活跃边界,扩展版本补充了完整证明和数值研究。
AI 中文摘要
我们研究增益调度线性二次调节的策略优化问题,其中通过固定加权函数插值得到的一组增益,将针对一组被控对象进行优化。所得代价会产生虚假局部极小值,而现有收敛性证明要么是局部的,要么极为保守。我们建立了一个精确恒等式:当用极小值点的闭环协方差计算代价梯度时,调度代价关于极小值点是星形凸的。该恒等式在整个可行集上成立,适用于调度的任意参数化形式。收敛性由一个无量纲比率决定:当该比率满足阈值条件时,梯度下降会以显式速率在整个次水平区域内线性收敛到最优解;而在每个虚假驻点处,该条件必然不成立。直接最大化该比率的实验表明,该阈值是代价景观的一个活跃边界。此扩展版本包含了因篇幅限制在原通讯中省略的完整证明和额外数值研究。
英文摘要
We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely conservative. We establish an exact identity: when the gradient of the cost is evaluated with the minimizer's closed-loop covariances, the scheduled cost is star-convex about the minimizer. The identity holds on the entire feasible set, for any parametrization of the schedule. Convergence is governed by a single dimensionless ratio. Wherever the ratio satisfies a threshold condition, gradient descent converges linearly to the optimum on entire sublevel regions at an explicit rate; at every spurious stationary point the condition necessarily fails. Experiments that maximize the ratio directly show the threshold to be an active boundary of the landscape. This extended version contains the complete proofs and additional numerical studies omitted from the letter for space.
Comments16 pages, 3 figures. Extended version of a manuscript submitted to IEEE Control Systems Letters. Code: https://github.com/Rainlabuw/hidden-star-convexity