arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36390stat.MEcs.LGmath.STstat.MLstat.TH

当动作是函数时拟合Q迭代的有限样本理论

Finite-Sample Theory for Fitted Q-Iteration When Actions Are Functions

  • The George Washington University(乔治华盛顿大学)
  • The University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Gefei Lin, Rui Miao, Xiaoke Zhang

AI总结:

针对动作是函数的离线强化学习,提出在评论家相对覆盖条件下研究平滑正则化策略搜索,为拟合Q迭代提供有限样本遗憾保证,并验证其有效性。

AI中文摘要:

离线强化学习旨在从先前收集的数据中寻求最优决策规则。在某些应用中,决策可以是整个函数,例如放射治疗中的通量图或机器人技术中的平滑运动轨迹。在本文中,我们研究了在折扣无限时域设置下具有函数动作的拟合Q迭代(FQI)的有限样本理论。在此设置中出现了三个主要困难:首先,函数动作缺乏勒贝格概率密度使覆盖描述复杂化;其次,常规覆盖要求可能具有限制性;第三,大的函数动作空间使得FQI中的贪心优化具有挑战性。为了解决这些困难,我们研究了在评论家相对覆盖条件下的平滑正则化策略搜索。该条件衡量记录数据如何区分相关的动作值差异,而不需要动作密度。我们的主要定理给出了相对于固定平滑函数动作策略类中最佳值的学习策略遗憾的有限样本保证。结果允许轨迹长度有界或增长,并且Q函数可以通过函数输入核岭回归或自适应函数神经网络来拟合。对于几个例子,我们可以获得在记录转换数量上多项式衰减的遗憾界,直至对数因子,具有对数多个FQI迭代。数值实验展示了学习到的函数动作策略相对于恒定动作策略的优势,并支持我们采用评论家相对覆盖条件。

英文摘要:

Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire function, such as a fluence map in radiation therapy or a smooth movement trajectory in robotics. In this paper, we study the finite-sample theory for fitted Q-iteration (FQI) with functional actions in a discounted infinite-horizon setting. Three major difficulties arise in this setting: first, the absence of a Lebesgue probability density for functional actions complicates coverage descriptions; second, conventional coverage requirements can be restrictive; and third, the large functional action space makes greedy optimization in FQI challenging. To address these difficulties, we study smoothness-regularized policy search under a critic-relative coverage condition. This condition measures how well logged data distinguish relevant action-value differences without requiring an action density. Our main theorem gives finite-sample guarantees for learned-policy regret relative to the best value within a fixed smooth class of functional-action policies. The results allow trajectory lengths to be either bounded or growing and Q-functions to be fitted by either functional-input kernel ridge regression or adaptive functional neural networks. For a few examples, we can obtain polynomially decaying regret bounds in the number of logged transitions, up to logarithmic factors, with logarithmically many FQI iterations. Numerical experiments show gains of learned functional-action policies over constant-action policies and support our adoption of a critic-relative coverage condition.

↑