arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

静态和时变网络上的方差减少Q学习

Variance-Reduced Q-Learning over Static and Time-Varying Networks

Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath

arXiv 2607.21876首次发表:更新:

发表机构

Department of Electrical and Computer Engineering, North Carolina State University; Department of Electrical and Computer Engineering, University of California San Diego(北卡罗来纳州立大学电气与计算机工程系; 加利福尼亚大学圣地亚哥分校电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多个智能体在同一MDP下的分散强化学习问题,提出基于轮次的分布式Q学习算法VRDQ,该算法在静态和时变网络中能实现高概率有限时间收敛,样本复杂度加速且通信量仅需\(\tilde{O}(1)\),改善了通信成本。

AI 中文摘要

我们研究了一个涉及多个与同一马尔可夫决策过程(MDP)交互的智能体的分散强化学习问题。智能体可通过网络交换信息以共同学习最优状态-动作值函数。为此,我们引入了一种新颖的基于轮次的分布式Q学习算法VRDQ,在每个轮次中,智能体局部估计贝尔曼最优算子并使用基于共识的协议传播信息。对于静态和时变网络,我们建立了VRDQ的高概率有限时间收敛率,其从协作中获得线性加速。关键的是,我们证明这种样本复杂度的加速仅需\(\tilde{O}(1)\)通信量,显著改善了先前工作中的通信成本。

英文摘要

We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed $Q$-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and diffuse information using a consensus-based protocol. For both static and time-varying networks, we establish high-probability finite-time convergence rates for VRDQ that enjoy linear speedups from collaboration. Crucially, we prove that such speedups in sample-complexity require only $\tilde{O}(1)$ communication, substantially improving upon the communication costs in prior work.

CommentsAccepted at the 2026 American Control Conference (ACC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑