发表机构
Department of Electrical and Computer Engineering, University of California, Riverside; Engineering Systems and Design Pillar, Singapore University of Technology and Design; Frontiers Science Center for Mobile Information Communication and Security, School of Mathematics, Southeast University(加州大学河滨分校电气与计算机工程系; 新加坡科技与设计大学工程系统与设计学院; 东南大学数学学院移动信息通信与安全前沿科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究网络多智能体强化学习,提出CDCPG算法。通过构造条件期望消除延续核不匹配,证明激励界,在特定条件下驱动平稳性度量,还有自适应局部规则。实验验证算法在局部性和特征维度方面的有效性。
AI 中文摘要
我们开发了连续分布式耦合策略梯度(CDCPG)算法,用于具有连续状态和动作空间的网络马尔可夫决策过程中的合作强化学习。每个智能体在有界图邻域上维护一个局部智能体,局部最小二乘时间差分评论家通过局部转移核的谱随机特征表示来评估截断动作值函数。分析有四个贡献。首先,截断动作值函数被构造为邻域上的条件期望,产生了一个适定的局部贝尔曼理论,消除了朴素截断论证中的延续核不匹配。其次,我们揭示了归一化随机特征的时间差分稳定性的维度障碍,并证明了一个无条件激励界,将稳定性降低到对称持续激励条件,可通过在线矩阵集中证书进行监测。第三,在智能体交互的指数空间衰减、激励条件和目标的光滑性下,CDCPG使用$\widetilde{\mathcal{O}}(\epsilon^{-2})$共享预言机样本将平均每个智能体的平稳性度量驱动到明确表征的近似下限的任何过剩$\epsilon$内,过剩依赖性与光滑非凸一阶速率匹配;每个智能体的计算和通信由邻域大小而非网络大小控制。第四,一个自适应局部规则选择平衡截断和图衰减残差与目标精度的半径。在网络线性二次基准上的实验证实了局部性和特征维度预测。
英文摘要
Learning local policies for continuous networked systems requires accounting for the effects of decisions beyond each agent's observation neighborhood. Spatial decay limits these effects, but a finite critic must also control representation and estimation errors throughout policy optimization. We analyze the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm using local random Fourier features and least-squares temporal-difference critics. For features that retain the boundary inputs required by the local dynamics, we derive an action-value representation with separate spatial and finite-feature residuals. A global integrated transition-approximation bound and a projected Bellman argument control population prediction error without an inverse-conditioning multiplier. We then quantify the dependence of critic estimation on feature excitation and dimension, and construct simultaneous lower confidence bounds for temporal-difference conditioning along the executed iterates. Combining critic error with localized reward aggregation bounds the expected squared projected-gradient mapping by an optimization term and an explicit residual separating spatial approximation, finite features, and omitted distant rewards. For fixed neighborhoods and feature dimension, the shared-oracle sample count is inverse-squared in the excess squared-stationarity accuracy, up to logarithmic factors. The guarantee assumes known local dynamics and rewards, independent discounted-occupancy samples, and stated excitation, decay, and smoothness conditions, and is conditional on favorable feature draws. Numerical studies illustrate related implementations on a linear-coupled-quadratic benchmark.
Commentsv2