发表机构
Yale University; London School of Economics and Political Science; George Washington University; University of Miami(耶鲁大学; 伦敦政治经济学院; 乔治华盛顿大学; 迈阿密大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出了稀疏加性离线策略评估框架,用稀疏加性结构的非线性函数建模Q函数,推导了与环境维度对数相关的误差界,还提出组稀疏性特征筛选方法,可在轨迹有限时实现准确的值函数估计。
AI 中文摘要
我们开发了一种新的框架,用于无限期强化学习的灵活、非线性且可解释的离线策略评估。为处理大状态空间并支持透明决策,我们采用具有稀疏加性结构的非线性函数类来建模Q函数。我们推导了估计目标策略值函数的高概率有限样本误差界,表明这些界仅与环境维度d呈对数相关,从而缓解了维度灾难。与大多数现有离线策略评估理论通常假设可获取大量轨迹不同,我们的分析保证当轨迹数量或时间范围足够大时,能实现准确的值函数估计。此外,我们提出了一种基于组稀疏性的特征筛选程序,该程序能以高概率识别出包含所有相关协变量的简化特征集。数值实验证明了所提方法的有效性。
英文摘要
We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probability finite-sample error bounds for estimating the value function of a target policy and show that the bounds depend only logarithmically on the ambient dimension $d$, thereby alleviating the curse of dimensionality. In contrast to most existing theory for off-policy evaluation, which typically assumes access to many trajectories, our analysis guarantees accurate value estimation when either the number of trajectories or the time horizon is sufficiently large. In addition, we propose a group-sparsity-based feature screening procedure that identifies, with high probability, a reduced feature set containing all relevant covariates. Numerical experiments demonstrate the effectiveness of the proposed approach.