带有乘性噪声的不定平均场社会优化的逆强化学习
Inverse reinforcement learning for indefinite mean-field social optimization with multiplicative noise
浏览论文内容
中文总结 AI 辅助
针对带有乘性噪声和不定代价权重的线性二次平均场社会优化问题,提出基于模型和无模型的逆强化学习算法,通过专家演示恢复未知社会代价权重,数值模拟验证了方法有效性。
中文摘要 AI 辅助
本文研究线性二次平均场(MF)社会优化的逆强化学习(RL)问题。所考虑的系统具有乘性噪声和不定代价权重,这违背了标准凸性假设并带来了分析挑战。目标是从专家演示中恢复未知的社会代价权重并复现最优控制策略,这需要求解耦合的随机代数Riccati方程和带有未知系统动力学的Lyapunov方程。为此,我们首先提出一种基于模型的逆RL算法,该算法包含两个顺序循环,分别处理个体和平均场动力学,我们证明了其收敛性和闭环可稳定性。此外,我们刻画了所恢复代价权重的非唯一性。为消除对系统动力学的依赖,我们开发了一种基于积分RL和最小二乘辨识的无模型逆RL算法,该算法仅需满足温和秩条件的实测轨迹数据。最后,数值模拟验证了所提方法的有效性。
英文摘要
This paper studies the inverse reinforcement learning (RL) problem for linear-quadratic mean-field (MF) social optimization. The considered system features multiplicative noise and indefinite cost weights, which violate standard convexity assumptions and pose analytical challenges. The goal is to recover unknown social cost weights from expert demonstrations and reproduce the optimal control policies. This requires solving coupled stochastic algebraic Riccati equations and Lyapunov equations with unknown system dynamics. To this end, we first propose a model-based inverse RL algorithm with two sequential loops that separately handle individual and MF dynamics, and we prove its convergence and closed-loop stabilizability. Moreover, we characterize the non-uniqueness of the recovered cost weights. To eliminate reliance on system dynamics, we develop a model-free inverse RL algorithm using integral RL and least-squares identification, which requires only measured trajectory data satisfying mild rank conditions. Finally, numerical simulations validate the effectiveness of the proposed approaches.