基于内在动机的探索驱动型个性化联邦强化学习
Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation
浏览论文内容
中文总结 AI 辅助
本研究提出EDPFRL-IM框架,将内在动机探索与随机网络蒸馏引入个性化联邦强化学习,在保护客户端隐私的同时,提升了延迟和稀疏奖励场景下的策略个性化与样本效率。
中文摘要 AI 辅助
个性化联邦强化学习(PFRL)采用去中心化方式存储和访问基于过往经验的信息,同时在学习每个客户端策略时保护其数据隐私。现有许多PFRL方法严重依赖利用现有强化学习奖励信号来推导每个客户端的最优策略,从而忽略了非平稳或稀疏奖励环境中的探索。本研究提出一种新的探索驱动型框架——基于内在动机的探索驱动型个性化联邦强化学习(EDPFRL-IM),该框架利用每个客户端固有的好奇心驱动型探索来促进局部探索并保护客户端隐私。此外,为了通过探索先前未探索的状态空间来促进策略发现,客户端将内在随机网络蒸馏(RND)信号添加到其外在奖励中。另外,服务器无法访问客户端的原始经验或局部梯度估计,而是发送全局探索先验并从每个客户端收集最少的新颖性摘要,以实现客户端之间多样化且协调的探索。在基准环境中的实验表明,我们的框架在策略个性化和样本效率方面优于平均PFRL基准,尤其在延迟和稀疏奖励系统中表现突出。总体而言,EDPFRL-IM能够将灵活的探索性学习结构集成到联邦强化学习系统中,同时保留客户端隐私。
英文摘要
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments. In this work, we introduce a new exploration-driven framework, Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), that leverages an inherent curiosity-driven exploration at each client to promote local exploration and protect client privacy. Furthermore, to facilitate policy discovery via exploration in previously unexplored state spaces, clients add an intrinsic random network distillation (RND) signal to their extrinsic reward. Additionally, the server does not have access to clients' raw experiences or local gradient estimates; instead, the server sends global exploration priors and collects minimal novelty summaries from each client to enable both diverse and coordinated exploration among clients. Experiments in benchmark environments show that our framework outperforms average PFRL benchmarks in policy personalization and sample efficiency, primarily in delayed and sparse reward systems. Overall, EDPFRL-IM enables the integration of a flexible exploratory learning structure into federated reinforcement learning systems while preserving client privacy.