arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有限群体LQG社会控制的分散策略:一种强化学习方法

Decentralized Strategies for Finite Population LQG Social Control: A Reinforcement Learning Approach

Liangyuan Guo, Bing-Chang Wang, Guangchen Wang

arXiv 2608.30244首次发表:更新:

AI 中文总结

本文针对带乘性噪声的有限群体LQG分散社会控制问题,提出一种无需系统矩阵先验知识的无模型强化学习算法,通过求解两个代数黎卡提方程实现社会最优,经数值示例验证有效。

AI 中文摘要

本文针对带有乘性噪声的有限群体线性二次高斯(LQG)分散社会控制问题,提出了一种新型无模型算法。代价函数中的状态和控制权重不局限于半正定。对于有限时域和无限时域两种情况,目标是通过求解两个代数黎卡提方程(ARE)得到社会最优解,且无需预先知晓系统矩阵。随后,完成了求解分散社会控制问题的无模型算法设计。特别地,在无限时域情况下,该算法的收敛性基于对李雅普诺夫型算子谱性质的分析,对比了有限时域与无限时域下强化学习(RL)解的差异。最后,通过数值示例验证了所提算法的有效性。

英文摘要

This paper presents a novel model-free algorithm for the finite-population linear quadratic Gaussian (LQG) decentralized social control problem with multiplicative noise. The state and control weights in the cost functional are not limited to be positive semidefinite. For both finite-horizon and infinite-horizon cases, the goal is to obtain a social optimum by solving two algebraic Riccati equations (AREs), without requiring prior knowledge of the system matrices. Then, we complete the design of a model-free algorithm for solving the decentralized social control problem. Especially, in the infinite-horizon case, the algorithm's convergence is based on analyzing the spectral property of the Lyapunov-type operator. The differences of reinforcement learning (RL) solutions between the finite-horizon and infinite-horizon cases are compared. Finally, the effectiveness of the proposed algorithm is demonstrated by a numerical example.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑