发表机构
California Institute of Technology; Massachusetts Institute of Technology; Purdue University; University of California, Riverside; University of Virginia; University of Illinois Urbana-Champaign(加州理工学院; 麻省理工学院; 普渡大学; 加州大学河滨分校; 弗吉尼亚大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对强化学习在线策略评估的高方差问题,提出双循环梯度算法学习鲁棒行为策略,其对转移扰动敏感性更低,可降低真实环境评估成本。
AI 中文摘要
在强化学习策略评估中,经典的同策略方法在估计策略性能时往往存在高方差问题。为缓解该问题,行为策略搜索被提出以学习专门用于降低在线评估方差的数据收集策略。然而,这些方法未考虑转移函数中的不确定性。在实际应用中,由于建模误差或近似限制,模拟器中的转移过程常与真实环境存在差异。因此,在仿真环境中训练得到的行为策略部署到真实环境时仍可能产生高方差,导致对真实世界评估样本的依赖成本高昂。本研究提出一种基于双循环梯度的算法,用于学习兼具高效性和对转移不确定性鲁棒性的行为策略。理论上,我们推导了新颖的转移方差梯度表达式,并为该算法建立了全局收敛保证。数值实验表明,与现有方法相比,我们的方法对转移扰动的敏感性更低,为其实用性提供了有力支持。
英文摘要
In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not account for uncertainties in the transition functions. In practice, simulator transitions often differ from the real world due to modeling errors or approximation limitations. As a result, behavior policies trained in simulation may still yield high variance when deployed in real environments, leading to costly reliance on real-world evaluation samples. In this work, we propose a double-loop gradient-based algorithm for learning behavior policies that are both efficient and robust to transition uncertainty. Theoretically, we derive novel transition-variance gradient expressions and establish global convergence guarantees for the algorithm. Numerically, we demonstrate that our method is less sensitive to transition perturbations than existing approaches, providing supportive evidence for its practical utility.