arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39981math.OC

分布LQR中随机回报的精确方差及其在均值-方差最优控制中的应用

ExactVariance of Random Return in Distributional LQR and Its Application to Mean-Variance Optimal Control

  • Imperial College London(帝国理工学院)
  • KTH Royal Institute of Technology(皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Ruyi Teng, Zifan Wang, Yulong Gao

AI总结:

针对经典LQR忽视性能变异性的缺陷,本文推导了分布LQR中随机回报方差的精确闭式解,并提出伴随梯度下降算法求解均值-方差最优控制问题,以显式权衡期望回报与风险。

AI中文摘要:

经典线性二次调节器(LQR)最小化期望累积回报,但未能考虑性能变异性,使其不足以应对风险敏感的应用。为解决此问题,我们将累积回报的方差作为LQR中的风险度量。我们在离散时间分布LQR框架内,针对具有对称概率密度的独立同分布扰动,推导了折扣无限期回报方差的第一个精确闭式表达式。对于高斯扰动,该表达式可简洁地简化为仅依赖于扰动协方差的形式。基于这些理论基础,我们构建了一个均值-方差最优控制问题,明确管理期望回报与性能变异性之间的权衡。为解决由此产生的非凸优化问题,我们针对原问题的惩罚形式提出了一种新颖的伴随梯度下降算法,并证明所有迭代保持稳定并收敛到惩罚目标的一个驻点。数值实验验证了该框架的有效性及固有的风险-性能权衡。

英文摘要:

The classical linear quadratic regulator (LQR) minimizes the expected cumulative return but fails to account for performance variability, rendering it inadequate for risk-aware applications. To address this, we introduce the variance of the cumulative return as a risk measure in LQR. We derive the first exact closed-form expression for the variance of the discounted in?finite horizon return within the discrete-time Distributional LQR framework, for i.i.d. disturbances with symmetric probability densities. For Gaussian disturbances, this expression elegantly simplifies to a form dependent only on the disturbance covariance. Leveraging these theoretical foundations, we formulate a mean-variance optimal control problem that explicitly manages the trade-off between expected return and performance variability. To address the resulting non-convex optimization problem, we propose a novel adjoint gradient descent algorithm for a penalized formulation of the original problem, and establish that all iterates remain stabilizing and converge to a stationary point of the penalized objective. The effectiveness of this framework and the inherent risk-performance trade-off? are demonstrated through numerical experiments.

补充信息

↑