发表机构
Carnegie Mellon Univeristy; Netflix; Cornell University; Stanford University; Rutgers University(卡内基梅隆大学; 网飞; 康奈尔大学; 斯坦福大学; 罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自适应采样下M-估计量泛函的推断问题,提出基于Neyman正交性的两种新方法,构造渐近有效置信区间,允许非参数估计干扰参数,并在模型误设定下保持有效,模拟验证了动态定价中的优势。
AI 中文摘要
强化学习和上下文赌博机算法在序贯决策应用中变得越来越普遍。当这些方法被部署在高风险领域时,人们不仅对学习有效策略越来越感兴趣,而且对在自适应数据收集下学习的量进行统计推断也越来越感兴趣。然而,在这些设置中天真应用的经典程序可能会失败:即使估计量是无偏的,它们的方差也会变得路径依赖,因此可能不是渐近正态的。越来越多的文献已经出现以改善这个问题,但解决方案往往是特定于问题的,并且通常依赖于工作模型的正确设定。在这项工作中,我们开发了一个统一框架,用于在自适应采样下构造渐近有效的置信区间,以覆盖非参数M-估计量的光滑泛函。在Neyman正交性下,我们提供了两种进行推断的新方法:(1)基于影响函数实现二次变差的自归一化统计量,和(2)使用基于重加权影响函数增量的条件方差插件估计的统计量。我们的结果允许对干扰参数进行灵活的非参数估计,并且在模型误设定下仍然有效。我们的理论得到了动态定价应用的模拟研究的支持,该研究表明,该方法可以产生渐近有效的置信区间,而标准方法在这些情况下会失败。
英文摘要
Reinforcement learning and contextual bandit algorithms have become increasingly common in sequential decision-making applications. When these methods are deployed in high-stakes domains, there is growing interest not only in learning effective policies, but also in conducting statistical inference for quantities learned under adaptive data collection. However, classical procedures applied naively in these settings can fail: even when estimators are unbiased, their variance becomes path-dependent and as a result may not be asymptotically normal. A growing literature has emerged to ameliorate this problem, but solutions tend to be problem specific and often rely on correct specification of a working model. In this work, we develop a unified framework for constructing asymptotically valid confidence intervals to cover smooth functionals of nonparametric M-estimands under adaptive sampling. Under Neyman orthogonality, we provide two novel methods for performing inference: (1) a self-normalized statistic based on the realized quadratic variation of the influence function and (2) a statistic using a plug-in estimate of the conditional variance based on reweighted influence function increments. Our results allow for flexible nonparametric estimation of nuisance parameters and remain valid under model misspecification. Our theory is supported by a simulation study for a dynamic pricing application which demonstrates that this method can produce asymptotically valid confidence intervals where standard methods fail.
Comments41 pages, 5 figures. Accepted at The Fortieth Annual Conference on Neural Information Processing Systems (NeurIPS 2026)