arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12753cs.LG

信息不对称下回合式马尔可夫决策过程中的去中心化多智能体Q学习

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

发表机构加州大学洛杉矶分校
查看机构详情
  • University of California, Los Angeles(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Larissa Xu, King Bi, William Chang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对三种信息不对称的回合式表格MDP去中心化多智能体强化学习问题,提出对应算法,证明其遗憾值在对数因子内匹配单智能体Q学习速率,边界适用于小智能体数量或小动作集场景。

中文摘要 AI 辅助

我们研究了回合式表格型马尔可夫决策过程(MDPs)中的去中心化多智能体强化学习,涉及三种信息不对称形式:(A)动作不可观测但奖励共享,(B)动作可观测但奖励独立,(C)动作不可观测且奖励独立。智能体在学习过程中无法通信,但可预先约定协议。针对问题A和B,我们提出了mQ-learning和mQ-learning-intervals算法,实现了$\tilde{O}(\boldsymbol{\text{sqrt}}(H^4 S A_{\text{joint}}\boldsymbol{\text{ }}T))$的遗憾值,其中$H$为回合长度,$S$为状态数量,$T = KH$为总步数,$A_{\text{joint}} = \boldsymbol{\text{prod}}_{i=1}^M |\boldsymbol{\text{mathcal{A}}}_i|$为$M$个智能体的联合动作空间。针对问题C,我们提出了mEXC和mEXC-Bellman两种两阶段“先探索后承诺”算法,遗憾值为$\tilde{O}(H (S A_{\text{joint}})^{1/3} T^{2/3})$。与中心化联合动作基准相比,信息不对称下的去中心化学习在对数因子范围内与Jin等人2018年提出的单智能体Q学习速率相当。由于$A_{\text{joint}}$随$M$呈指数增长,该边界对小$M$或单智能体小动作集的场景最具意义。

英文摘要

We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards. Players cannot communicate during learning but may agree on a protocol a priori. For Problems A and B we propose \texttt{mQ-learning} and \texttt{mQ-learning-intervals}, achieving $\tilde{O}(\sqrt{H^4 S A_{\text{joint}}\, T})$ regret, where $H$ is the horizon, $S$ the state count, $T = KH$ the total steps, and $A_{\text{joint}} = \prod_{i=1}^M |\mathcal{A}_i|$ the joint action space across $M$ players. For Problem C we give \texttt{mEXC} and \texttt{mEXC-Bellman}, two-phase explore-then-commit algorithms with regret $\tilde{O}(H (S A_{\text{joint}})^{1/3} T^{2/3})$. Against the centralized joint-action benchmark, decentralized learning under information asymmetry matches the single-agent Q-learning rate of \cite{jin2018q} up to logarithmic factors. Because $A_{\text{joint}}$ grows exponentially in $M$, the bounds are most meaningful for small $M$ or small per-player action sets.

↑