arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

论策略优化中成功条件化的收敛性

On the Convergence of Success Conditioning for Policy Optimization

Matthew Brun, Xu Andy Sun

arXiv 2610.03642首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究成功条件化策略在随机环境中的收敛性,证明其在广泛MDP类上收敛到最优策略,并给出折扣和单周期情形下的收敛速率。

AI 中文摘要

成功条件化是一种在随机环境中改进决策策略的策略;它通过增加产生成功结果的动作的概率来更新策略。成功条件化常见于许多强化学习应用中,但其极限行为和收敛速率尚未被充分理解。在这项工作中,我们证明了成功条件化在一类广泛的马尔可夫决策过程(MDPs)上收敛到最优策略。我们还推导了在某些常见设置下的收敛速率。对于折扣MDP,我们证明了在$\mathcal{O}(1/\varepsilon^p)$次迭代内收敛到$\varepsilon$-最优策略,其中指数$p$取决于问题数据。对于单周期MDP,这样的策略在$\mathcal{O}(\log(1/\varepsilon))$次迭代内获得。

英文摘要

Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, such a policy is obtained within $\mathcal{O}(\log(1/\varepsilon))$ iterations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑