arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09981stat.MLcs.LGstat.ME

强化学习中的最优值推断

Optimal Value Inference for Reinforcement Learning

Nan Lu, Ethan Lee, James M. Robins, David Simchi-Levi, Junwei Lu

AI总结:

本文研究强化学习中最优值的离线推断,通过自诱导贝尔曼方程推导新干扰参数,结合Neyman正交性提出去偏估计器,实现发散时间范围下的有效推断,并验证于合成实验及实际决策问题。

AI中文摘要:

我们研究强化学习中最优值的离线推断。我们推导出两个新的干扰参数,它们作为自诱导贝尔曼方程的不动点,其中我们通过其softmax对应来近似最大贝尔曼算子。我们通过Neyman正交性提出一个去偏估计器,并在发散的时间范围下建立其渐近正态性,即使行为策略随时间变化,只要干扰参数具有许多机器学习方法可达到的统计速率。我们为这些干扰参数提供了具体的估计程序,并表明它们可以导致有效的推断。合成实验验证了我们推断方法的数值性能,并将其应用于现实生活中的决策问题,包括自行车重新定位和AI智能体工具使用。

英文摘要:

We study offline inference for the optimal value in reinforcement learning under finite state and action spaces. Two new nuisances are derived as fixed points of a self-induced Bellman equation, in which we approximate the maximum Bellman operator by its softmax correspondence. We propose a debiased estimator through the Neyman orthogonality and establish its asymptotic normality under diverging horizons even when the behavior policy changes with time, as long as the nuisances have the statistical rates that can be achieved by many machine learning methods. We provide a concrete estimating procedure for these nuisances and show they can lead to valid inference. Synthetic experiments validate the numerical performance of our inference method, and we implement it in real-life decision-making problems, including bike repositioning and AI agentic tool use.

↑