具有线性函数近似和对数通信成本的可证明高效联邦强化学习
Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost
浏览论文内容
中文总结 AI 辅助
针对联邦强化学习的线性函数近似场景,提出Fed-LSVI算法,通过压缩统计量实现对数通信成本,达到最优遗憾界,解决了通信成本高与隐私约束的问题。
中文摘要 AI 辅助
我们研究了带有线性函数近似的联邦在线强化学习。虽然近期多智能体强化学习算法能提供强遗憾保证,但它们通常需要共享原始轨迹,这种依赖会产生与回合数成线性比例的通信成本,且违反联邦设置的隐私约束。为解决这些局限,我们提出Fed-LSVI,这是首个针对回合马尔可夫决策过程中带有线性函数近似的在线强化学习的可证明高效联邦算法。通过整合基于行列式的事件触发同步与逐步反向更新机制,Fed-LSVI使智能体仅通过交换压缩的充分统计量即可协作学习最优策略。我们证明Fed-LSVI实现了$\tilde{\bigO}(\boldsymbol{\textit{Md}^3\textit{H}^4\textit{T}})$的遗憾界,其中$d$为特征维度,$H$为回合长度,$M$为智能体数量,$T$为每个智能体的回合数,这与带有线性函数近似的多智能体在线强化学习的最优已知遗憾相匹配。此外,通过遵循联邦设置的严格通信和隐私约束,Fed-LSVI将通信成本降低至仅与$T$呈对数依赖,较现有方法有显著改进。
英文摘要
We study federated online reinforcement learning with linear function approximation. While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, they typically require sharing raw trajectories. This reliance incurs a communication cost that scales linearly with the number of episodes and violates the privacy constraints of federated settings. To address these limitations, we propose Fed-LSVI, the first provably efficient federated algorithm for online reinforcement learning with linear function approximation in episodic Markov decision processes. By integrating a determinant-based event-triggered synchronization with a stepwise backward update mechanism, Fed-LSVI enables agents to collaboratively learn an optimal policy by exchanging only compressed sufficient statistics. We prove that Fed-LSVI achieves a regret bound of $\widetilde{\mathcal O}(\sqrt{Md^3H^4T})$, where $d$ is the feature dimension, $H$ is the horizon length, $M$ is the number of agents, and $T$ is the number of episodes per agent, matching the best-known regret for multi-agent online reinforcement learning with linear function approximation. Moreover, by following the stringent communication and privacy constraints of the federated setting, Fed-LSVI reduces the communication cost to only logarithmic dependence on $T$, representing a significant improvement over prior methods.
发表机构
- The Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。