arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06467cs.LG

执行强化学习中的局部与全局稳定性

Local and Global Stability in Performative Reinforcement Learning

  • University of Warwick(华威大学)

机构由 AI 辅助整理,请以论文原文为准。

Debmalya Mandal

AI总结:

本文研究执行强化学习中策略混合的稳定性,提出局部与全局稳定性概念,证明加权Hedge算法在无敏感性假设下实现局部稳定,并引入有界转移范围假设保证全局稳定,扩展到多智能体博弈。

AI中文摘要:

在执行强化学习中,部署的策略会塑造生成学习者未来数据的环境,其自然解概念是执行稳定策略,即在该策略诱导的环境中为最优策略。现有的收敛保证依赖于环境映射 π ↦ (P_π, r_π) 的 Lipschitz 敏感性假设,这些假设难以验证,并且在多智能体最佳响应动态等场景中失效。我们转而研究策略混合的稳定性,并表明由此产生的图景与执行预测根本不同,在执行预测中,随机化消除了对任何敏感性假设的需求。我们区分局部混合稳定性(一种占用加权的二阶松弛,我们证明其等价于平稳性)和全局混合稳定性(其针对任意偏离策略提供保证)。我们的第一个结果是,加权逐状态 Hedge 动态以 O(1/√T) 的速率将局部稳定性差距降至零,适用于任意可能不连续的环境映射,无论是精确反馈还是轨迹反馈。这两个概念确实不同:我们展示了一个实例,其中局部稳定性精确实现,但每个混合的全局稳定性差距都远离零。对于全局稳定性,我们引入了有界转移范围假设,该假设严格弱于 Lipschitz 敏感性,在此假设下,未加权的逐状态 Hedge 收敛到 O(γ ε_P/(1-γ)^3) 的下限,并且我们证明了在轨迹反馈下匹配 ε_P 的 Ω(γ ε_P/(1-γ)) 下界,因此该下限是不可避免的。最后,我们将这两个概念扩展到 n 人执行马尔可夫博弈,在联合环境映射或博弈结构无任何假设的情况下获得局部稳定性,并为执行马尔可夫势博弈获得全局稳定性。

英文摘要:

In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $π\mapsto (P_π, r_π)$, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for mixtures of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish local mixed stability, an occupancy-weighted first-order relaxation that we show is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies. Our first result is that a weighted per-state Hedge dynamic drives the local stability gap to zero at an $O(1/\sqrt{T})$ rate for an arbitrary, possibly discontinuous, environment map, both with exact and with trajectory feedback. The two notions genuinely differ: we exhibit an instance where local stability is achieved exactly but every mixture has global stability gap bounded away from zero. For global stability we introduce a bounded transition range assumption, strictly weaker than Lipschitz sensitivity, under which unweighted per-state Hedge converges up to a floor of $O(γε_P/(1-γ)^3)$, and we prove a matching-in-$ε_P$ lower bound of $Ω(γε_P/(1-γ))$ under trajectory feedback, so this floor is unavoidable. Finally, we extend both notions to $n$-player performative Markov games, obtaining local stability with no assumption on the joint environment map or game structure, and global stability for performative Markov potential games.

↑