arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从不可靠轨迹中学习:对抗鲁棒的联邦Q学习

Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning

Sreejeet Maity, Aritra Mitra

arXiv 2610.06918首次发表:更新:

发表机构

North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Robust Async-Fed-Q算法,结合方差缩减与鲁棒聚合,解决联邦Q学习中对抗性智能体损坏信息的问题,在保持协作收益的同时实现对抗鲁棒性,并给出匹配的上下界及更优通信复杂度。

AI 中文摘要

我们研究联邦强化学习,其中多个智能体与一个共同的马尔可夫决策过程交互,并通过中央服务器通信,以协作学习最优状态-动作值函数。我们的目标是理解当一部分智能体表现对抗性并传输任意损坏的信息时,协作的样本效率优势是否能够保留。为解决此问题,我们引入Robust Async-Fed-Q,一种基于epoch的联邦学习算法,该算法在智能体端结合了Bellman最优性算子的方差缩减估计,并在服务器端进行鲁棒聚合。我们建立了高概率有限时间保证,表明所提方法在容忍对抗性损坏的同时,保留了诚实智能体间协作的统计增益。特别地,对抗性智能体的影响随着每个诚实智能体收集的数据量增加而减小,并在无限样本极限中最终消失。我们用信息论下界补充这些保证,该下界刻画了对抗性损坏不可避免的统计成本,从而为对抗鲁棒联邦强化学习提供了首个几乎匹配的上下界。我们进一步扩展框架以适应单轨迹马尔可夫采样和异构部分覆盖,其中不同智能体可能探索状态-动作空间的不同区域,学习依赖于它们的集体覆盖。最后,我们的基于epoch的设计显著改善了异步采样下联邦Q学习的最佳已知通信复杂度。

英文摘要

We study federated reinforcement learning in which multiple agents interact with a common Markov decision process and communicate through a central server to collaboratively learn the optimal state-action value function. Our goal is to understand whether the sample-efficiency benefits of collaboration can be retained when a fraction of the agents behave adversarially and transmit arbitrarily corrupted information. To address this problem, we introduce Robust Async-Fed-Q, an epoch-based federated learning algorithm that combines variance-reduced estimation of the Bellman optimality operator at the agents with robust aggregation at the server. We establish high-probability finite-time guarantees showing that the proposed method preserves the statistical gains of collaboration among the honest agents while tolerating adversarial corruption. In particular, the effect of the adversarial agents decreases as the amount of data collected by each honest agent grows and eventually vanishes in the infinite-sample limit. We complement these guarantees with information-theoretic lower bounds that characterize the unavoidable statistical cost of adversarial corruption, leading to the first nearly matching upper and lower bounds for adversarially robust federated reinforcement learning. We further extend our framework to accommodate single-trajectory Markovian sampling and heterogeneous partial coverage, where different agents may explore different regions of the state-action space and learning relies on their collective coverage. Finally, our epoch-based design substantially improves the best known communication complexity for federated Q-learning under asynchronous sampling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑