PRIME:无人机辅助应急通信网络多智能体环境中的可塑性恢复
PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks
- Department of Information and Communication Engineering, Kitami Institute of Technology, Japan(信息与通信工程系,函馆技术学院,日本)
- Graduate School of Informatics and Engineering, University of Electro-Communications, Japan(信息与工程研究生院,电子通信大学,日本)
- School of Computer Science and Technology, Anhui University of Technology, China(计算机科学与技术学院,安徽理工大学,中国)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究无人机辅助应急通信网络多智能体环境中因非平稳性致神经元休眠影响学习的问题,提出PRIME方法,扩展双向沉默神经元框架,聚合统计信息验证后干预,实验显示该方法提升回报并降低休眠神经元比例,明确扰动成本与沉默子空间维度有关。
AI中文摘要:
大多数用于这些网络的强化学习控制器假定条件是静止的,少数处理变化的控制器只对外界环境做出反应,而不检查网络的内部状态。我们表明持续的非平稳性会直接损害这种内部状态:随着目标的转移,神经元逐渐休眠,共享策略失去学习能力。在共享参数多智能体训练下,重置休眠神经元这种明显的补救方法是不安全的。PRIME(多智能体环境中的可塑性恢复)因此在干预前会对两个方向进行验证。它将双向沉默神经元框架扩展到合作多智能体强化学习,在整个团队批次上聚合激活和梯度统计信息,从训练损失已经沉积的梯度中读取反向信号,而不是从手工制作的代理中读取,并仅重新初始化同时处于激活休眠和梯度沉默状态的神经元。在学习能力恢复的同时保留了有用的表示。在一个相位切换无人机应急通信模拟器上,PRIME比MAPPO将四分位间距平均回报提高了24.9%,并将休眠神经元比例保持在10%-20%,而MAPPO为40%-45%;消融实验将收益归因于梯度信号和团队级聚合,而不是特定的重置算子。一个动态遗憾界表明,扰动成本与小的沉默子空间维度成比例,而不是与整个参数数量成比例。
英文摘要:
Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn. The obvious remedy, resetting dormant neurons, is unsafe under shared-parameter multi-agent training: many neurons that appear inactive are still receiving strong training gradients, and whether a neuron appears dormant depends on which agent's observations it processes. PRIME (Plasticity Recovery In Multi-agent Environments) therefore verifies both directions before intervening. Extending the bidirectional Silent Neuron framework to cooperative multi-agent reinforcement learning, it aggregates activation and gradient statistics over the full team batch, reads the backward signal from the gradient the training loss has already deposited , not from a hand-crafted proxy, and reinitializes only neurons that are simultaneously activation-dormant and gradient-silent. Useful representations are preserved while learning capacity is restored. On a phase-switching UAV emergency communication simulator, PRIME improves interquartile mean return by 24.9\% over MAPPO and holds dormant neuron fractions at 10--20\% versus 40--45\%; ablations attribute the gains to the gradient signal and team-level aggregation rather than to the specific reset operator. A dynamic regret bound shows that the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.