发表机构
State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制科学与工业技术国家级重点实验室、 cyber-系统与控制中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种无评论家的策略迭代方法,通过策略空间Riccati方程直接求解连续时间零和博弈的鞍点策略,并利用数据驱动算法实现未知动力学下的策略恢复,显著降低计算与内存开销。
AI 中文摘要
本文针对连续时间线性零和博弈提出了一种无评论家策略迭代(PI)方法。核心思想是直接在联合策略空间中刻画鞍点策略,而非将二次值矩阵视为迭代变量。引入了一个策略博弈Riccati方程(PGRE),其未知量仅为策略增益。证明了其解与博弈代数Riccati方程的对称解一一对应。基于一个稳定锚点,直接在actor空间中进行PI。在每一个稳定策略处,actor空间的雅可比矩阵非奇异,且所得策略序列与同步PI一致。对于未知动力学,一种数据驱动算法使用单批数据和端点增量的零空间投影来消除值矩阵,从而得到仅含actor的回归。建立了唯一策略恢复的充要秩条件,并证明该条件在迭代中不变,表明评论家可辨识性并非必要。一个电力系统频率调节实例验证了收敛性和策略恢复,而可扩展性测试表明计算和内存需求大幅降低。
英文摘要
This paper develops a critic-free policy iteration (PI) method for continuous-time linear zero-sum games. The central idea is to characterize the saddle-point policies directly in the joint policy space, rather than treating the quadratic value matrix as an iterative variable. A policy game Riccati equation (PGRE) is introduced whose unknowns are policy gains only. Its solutions are shown to be in one-to-one correspondence with the symmetric solutions of the game algebraic Riccati equation. Based on a stabilizing anchor, PI is performed directly in the actor space. The actor-space Jacobian is nonsingular at every stabilizing policy, and the resulting policy sequence coincides with that of simultaneous PI. For unknown dynamics, a data-driven algorithm uses a single batch of data and nullspace projection of endpoint increments to eliminate the value matrix, yielding an actor-only regression. A necessary and sufficient rank condition for unique policy recovery is established and shown to be iteration-invariant, demonstrating that critic identifiability is unnecessary. A power systems frequency-regulation example verifies convergence and policy recovery, while scalability tests demonstrate substantial reductions in computational and memory requirements.