arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散环境中多臂老虎机的策略梯度算法:收敛性与悔憾分析

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

Yanwei Jia, Du Ouyang

arXiv 2607.29593首次发表:更新:

发表机构

The Chinese University of Hong Kong; Tsinghua University(香港中文大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对扩散环境中多臂老虎机的策略梯度算法,证明其收敛性并推导O(log T)阶悔憾上界,改进了现有分析且可推广至离散时间算法。

AI 中文摘要

本文研究由Wang等人(2020)、Jia和Zhou(2022b)提出的连续时间强化学习框架下,由随机微分方程(SDE)描述的扩散环境中多臂老虎机问题的策略梯度更新。采用随机策略的logit参数化,证明其在任意常数学习率下几乎必然收敛到最优臂;推导了常数学习率低于时不变阈值时的非渐近悔憾上界,其阶为O(log T)。通过构造新型李雅普诺夫函数,改进了Lattimore(2026a)对同一SDE的分析,并展示了利用SDE工具分析策略梯度的透明性;此外,该李雅普诺夫函数也有助于分析离散时间策略梯度算法。

英文摘要

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the constant learning rate is below a time-invariant threshold; and the regret bound has order $O(\log T)$. We improve the analysis in Lattimore (2026a) for the same SDE by constructing a novel Lyapunov function and demonstrate the transparency of analyzing policy gradient using the tools in SDEs. In addition, the same Lyapunov function is also helpful in analyzing the discrete-time policy gradient algorithm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑