arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带折扣的指数效用强化学习的有限时间分析

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Ankur Naskar, Vivek T A, Aditya Kumar, Gugan Thoppe, Prashanth L. A

arXiv 2608.01917首次发表:更新:

AI 中文总结

本研究针对带折扣的指数效用强化学习,为两种无模型不动点算法建立了异步马尔可夫采样下的有限时间收敛速率,采用无参数步长,提供了首个有限时间保证。

AI 中文摘要

带折扣的指数效用为风险敏感的序列决策提供了有原则的准则,但其非线性结构使强化学习变得复杂。近期一项研究[Thoppe等人,2026]通过引入与贝尔曼兼容的替代函数,以及两种无模型的不动点算法来优化平稳策略下的该替代函数,解决了这一难题。不过,他们的主要收敛结果是渐近性的。在本研究中,我们在异步马尔可夫采样下,为上述两种算法建立了迭代次数为n时的有限时间速率$\tilde{O}(1/\tilde{n})$,其中$\tilde{O}$隐藏了对数表达式。重要的是,我们采用无参数的步长参数选择来推导这些速率结果。对于算法更简单的单时间尺度方法,其主要挑战在于更新方程与其底层幂律算子的收缩几何结构不直接匹配。我们通过利用算子的有界性、单调性和齐次性,为相对误差动力学获得局部伪收缩特性,克服了这种不匹配。随后,我们使用基于Moreau包络的李雅普诺夫函数和Polyak-Ruppert平均,在无参数步长下获得所述收敛速率。对于双时间尺度方法,主要挑战在于控制较快时间尺度上的跟踪误差。这些结果为无模型带折扣的指数效用强化学习提供了首个有限时间保证。

英文摘要

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies. However, their main convergence results are asymptotic. In this work, we establish finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions. Importantly, we employ parameter-free choices for the stepsize parameter to derive these rate results. For the algorithmically simpler one-timescale method, the main challenge is that its update equation is not directly aligned with the contraction geometry of its underlying power-law operator. We overcome this mismatch by exploiting the boundedness, monotonicity, and homogeneity of the operator to obtain a local pseudo-contraction property for the relative-error dynamics. We then use a Moreau-envelope-based Lyapunov function and Polyak--Ruppert averaging to obtain the stated convergence rate with parameter-free stepsizes. For the two-timescale method, the main challenge is to control a tracking error on the faster timescale. These results provide the first finite-time guarantees for model-free discounted exponential-utility reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑