arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连续动作赌博机的统计推断

Statistical Inference for Continuous Action Bandits

Raphael C Kim, Michele Santacatterina, Ramin Zabih, Rajarshi Mukherjee, Ivan Diaz

arXiv 2610.09801首次发表:更新:

发表机构

Cornell Tech, Cornell University; New York University School of Medicine; T.H. Chan School of Public Health(康奈尔大学康奈尔科技校区; 纽约大学医学院; 陈曾熙公共卫生学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对连续动作自适应收集数据,扩展核平滑双稳健估计,研究自适应加权估计量的性质,给出遗憾下界,并提出实验设计建议,通过模拟验证。

AI 中文摘要

在自适应收集数据下的统计推断在电子商务和移动健康领域日益流行。针对离散动作设定,已有多种方法,从加权到去偏方法。尽管连续动作在从最优定价到精准给药等实验中普遍存在,但据我们所知,支持连续动作自适应收集数据下统计推断的方法仍不完善。在本工作中,我们将核平滑双稳健估计从独立同分布数据扩展到具有连续动作的自适应收集数据。我们研究了一族自适应加权估计量,建立了其均方误差率、渐近正态性,并刻画了作为遗憾函数的估计下界。最后,我们提出了关于实验设计的建议,以兼顾效率与遗憾,并通过广泛的模拟研究支持我们的理论。

英文摘要

Statistical inference under adaptively collected data is becoming increasingly popular across e-commerce and mobile health. Many methods exist for discrete action settings, ranging from weighting to debiasing approaches. Despite the ubiquity of continuous actions in experimentation from optimal pricing to precision dosing, to the best of our knowledge, methods that support statistical inference under continuous action, adaptively collected data remain underdeveloped. In this work, we extend kernel-smoothed doubly robust estimation from i.i.d. data to adaptively collected data with continuous actions. We study a family of adaptively weighted estimators, establish their mean squared error rates, asymptotic normality, and characterize an estimation lower bound as a function of regret. We conclude with recommendations for experimental design informing efficiency and regret, and support our theory with an extensive simulation study.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑