发表机构
University of Illinois Urbana-Champaign; The University of Texas at Austin; The Chinese University of Hong Kong, Shenzhen(伊利诺伊大学厄巴纳-香槟分校; 德克萨斯大学奥斯汀分校; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对熵正则化线性二次控制问题,将Wasserstein策略梯度简化为有限维ODE并证明其指数收敛,明确了收敛指数在熵温度趋于0时的极限特性。
AI 中文摘要
Wasserstein策略梯度(WPG)通过在动作空间中进行传输来更新状态条件动作律。我们研究熵正则化的折扣线性二次(LQ)控制。贝尔曼验证论证表明,无约束问题具有线性高斯最优策略,且折扣占用加权的逐状态Wasserstein梯度与该策略类相切。因此,WPG可精确简化为反馈增益和动作协方差的有限维常微分方程(ODE)。我们证明该ODE整体适定,且从任意可允许初始化出发指数收敛。对于每个固定的LQ问题,当熵温度τ趋于0时,收敛指数存在正极限,其不含形如exp(-c/τ)的微扰因子,同时保留对控制问题条件数的常规依赖关系。
英文摘要
Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form $\exp(-c/τ)$, while retaining the usual dependence on the conditioning of the control problem.