发表机构
Stockholm University; Mahidol University(斯德哥尔摩大学; 玛希隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Per-FedAvg-PG算法,通过MAML式策略初始化实现个性化联邦强化学习,证明其收敛性并分析免Hessian变体,实验表明以更低样本成本实现有效迁移。
AI 中文摘要
我们研究个性化联邦强化学习,其中$n$个智能体各自在其马尔可夫决策过程中行动,通过服务器协作学习一个共享的MAML式策略初始化,该初始化在单个智能体通过一次局部策略梯度步骤进行适配后对其有效。我们提出Per-FedAvg-PG,其中智能体在通信轮次之间采取$\tau$步局部随机元策略梯度步骤,并证明其在$K=\mathcal O(\varepsilon^{-3/2})$轮次内达到个性化目标函数的$\varepsilon$-近似一阶驻点,且局部步骤数为$\tau=\Theta(\varepsilon^{-1/2})$。该分析依赖于强化学习设置的一个结构特征:在标准策略类正则性下,每个智能体的目标函数具有一致有界的梯度和Hessian,且带有显式常数,因此监督学习理论所施加的有界梯度和有界异质性条件自动满足,无需单独的异质性假设。精确元梯度需要内循环策略Hessian,我们的实验将其识别为实际瓶颈。因此,我们分析免Hessian变体,界定其偏差,并展示元梯度非零且数量级为$\alpha$的固定点,表明由此产生的驻点下限是方法本身的属性而非界限的属性。在表格和神经导航上的实验证实了预测行为,并展示了在比独立训练低一个数量级的样本成本下对未见智能体的迁移。这些结果共同将适配步长识别为可调个性化旋钮,将曲率估计识别为决定精确元梯度是否可行的关键量。
英文摘要
We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step. We propose Per-FedAvg-PG, in which agents take $τ$ local stochastic meta-policy-gradient steps between communication rounds, and prove that it reaches an $\varepsilon$-approximate first-order stationary point of the personalized objective in $K=\mathcal O(\varepsilon^{-3/2})$ rounds with $τ=Θ(\varepsilon^{-1/2})$ local steps. The analysis rests on a structural feature of the reinforcement learning setting: under standard policy-class regularity, the per-agent objectives have uniformly bounded gradients and Hessians with explicit constants, so the bounded-gradient and bounded-heterogeneity conditions imposed by the supervised theory hold automatically and no separate heterogeneity assumption is needed. The exact meta-gradient requires the inner-loop policy Hessian, which our experiments identify as the practical bottleneck. We therefore analyze the Hessian-free variant, bound its bias, and exhibit fixed points at which the meta-gradient is nonzero and of order $α$, showing that the resulting stationarity floor is a property of the method rather than of the bound. Experiments on tabular and neural navigation confirm the predicted behavior and show transfer to unseen agents at an order of magnitude lower sample cost than independent training. Together these results identify the adaptation step size as a tunable personalization knob and the curvature estimate as the quantity that governs whether exact meta-gradients are affordable.
Comments44 pages, 8 figures, Code: https://github.com/AliBeikmohammadi/Per-FedAvg-PG