发表机构
North Carolina State University(北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对含对抗智能体的联邦强化学习,提出Robust Fed-Q算法,融合模型基与模型无关方法及中位数均值技术,实现近乎零通信下的精确收敛与近最优有限时间速率。
AI 中文摘要
我们考虑一个联邦强化学习设置,其中包含$M$个智能体,所有智能体都与一个共同的马尔可夫决策过程(MDP)交互。智能体通过中央服务器交换信息以学习最优价值函数。我们的目标是理解在此类设置中,当一小部分智能体具有对抗性且可任意行动时,协作样本复杂度加速能达到何种程度。为此,我们提出Robust Fed-Q,一种联邦Q学习算法,它融合了基于模型和无模型强化学习的思想,以及鲁棒统计中的中位数均值设备。我们证明,尽管存在腐败,Robust Fed-Q以高概率(i)在无限样本极限下保证精确收敛到最优价值函数,并且(ii)享有近乎最优的有限时间速率,该速率受益于协作。此外,我们的方法仅需$\ ilde{O}(1)$轮通信即可实现上述每项保证,这一特性在通信是主要瓶颈的联邦学习中具有独立意义。
英文摘要
We consider a federated reinforcement learning setting involving $M$ agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q}, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just $\tilde{O}(1)$ rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.
CommentsAccepted at the 2026 American Control Conference (ACC 2026)
Journal refS. Maity and A. Mitra, "Robust Federated Q-Learning with Almost No Communication," 2026 American Control Conference (ACC), New Orleans, LA, USA, 2026, pp. 462-469