网络上的完全分布式与安全感知多智能体强化学习控制
Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks
浏览论文内容
中文总结 AI 辅助
本文提出一种完全分布式且安全感知的多智能体强化学习算法,通过图编码器与状态估计器增强系统感知,并利用控制屏障启发式减少节点忽视,在持续监测等网络控制问题上实现比集中式策略低26.3%的平均不确定性。
中文摘要 AI 辅助
本文开发了一种安全且完全分布式的多智能体强化学习(MARL)算法,用于解决网络上的一类离散时间控制问题,包括持续监测问题。智能体的完全分布式控制虽然提供了诸多好处,但面临指数级增长的样本复杂度、缺乏系统全局信息以及智能体间协调困难等问题。为解决这些问题,本文引入了一种完全分布式的多智能体强化学习算法,将深度强化学习与安全考量相结合。该方法将网络状态的局部观测历史输入两个并行的神经网络分支:图编码器(添加节点间的结构信息和相关性)和状态估计器(预测图中每个节点的不确定性)。此外,将该输入输入到演员-评论家网络的结果通过离散时间控制屏障启发式方法,以减少任何节点被忽视的可能性。这种方法通过增强系统感知能力并内置安全措施以防止采用潜在有害的控制策略,使完全分布式的智能体团队能够解决具有挑战性的问题。来自自定义模拟环境的数值结果表明,所提出的算法相比集中式控制策略实现了26.3%更低的平均不确定性,并且在不确定性性能上与添加了注意力层的计算复杂度更高的算法相差在1%以内。
英文摘要
This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning with safety considerations. This method feeds a history of local observations of the network's state into two parallel neural-network branches: the graph encoder, which adds structural information and correlations among nodes, and a state estimator, which predicts the uncertainty at each node in the graph. Additionally, the result of feeding that input into an actor-critic network is passed through a discrete-time control barrier heuristic to reduce the likelihood that any node will be neglected. This approach enables teams of fully decentralized agents to solve challenging problems by increasing system awareness and incorporating built-in safety measures to prevent the adoption of potentially harmful control policies. Numerical results from a custom simulation environment demonstrate that the proposed algorithm achieves 26.3 percent lower average uncertainty than a centralized control policy and is within 1 percent of the uncertainty performance of a more computationally complex algorithm with added attention layers.
发表机构
- Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。