发表机构
Yale University; Harvard University; Dartmouth College; University of Florida(耶鲁大学; 哈佛大学; 达特茅斯学院; 佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对时空干扰下的动态策略评估与学习,提出半参数加性模型下的稳定估计量族及数据自适应最优选择,并用于伊拉克援助分配以减少叛乱袭击。
AI 中文摘要
尽管序列决策在各个领域无处不在,但由于空间溢出效应和时间延续效应,利用时空数据进行策略评估与学习仍然具有挑战性。我们开发了在时空干扰下评估和学习个体化动态策略的方法。在允许复杂溢出和延续效应的半参数加性结果模型下,我们考虑一族稳定估计量来评估给定个体化策略的性能。从这一族中,我们选择一个数据自适应的最优估计量,以最小化渐近方差。随后,我们推导了所提出的策略评估估计量的渐近分布,并建立了我们的策略学习估计量的有限样本遗憾界。我们进一步提出了一种统计检验,通过确定交互作用的适当阶数来选择半参数加性模型的复杂度。通过模拟,我们评估了估计量的有限样本性能以及所提出检验的有效性。我们的动机性应用考察了2007年2月至2008年7月期间伊拉克经济援助的最优分配。利用解密的冲突数据,我们研究了援助项目在各地区间的每周分配,并发现将援助从持续高暴力水平的地区重新分配出去可以大幅减少叛乱袭击。
英文摘要
Although sequential decision-making is ubiquitous across domains, policy evaluation and learning with spatio-temporal data remain challenging due to spatial spillover and temporal carryover effects. We develop methods for evaluating and learning individualized dynamic policies under spatio-temporal interference. Under a semiparametric additive outcome model that allows for complex spillover and carryover effects, we consider a family of stabilized estimators for evaluating the performance of a given individualized policy. From this family, we select a data-adaptive optimal estimator that minimizes the asymptotic variance. We then derive the asymptotic distribution of the proposed policy evaluation estimator, and establish the finite-sample regret bounds of our policy learning estimator. We further propose a statistical test to select the complexity of the semiparametric additive model by determining the appropriate order of interactions. Through simulations, we assess the finite-sample performance of our estimators and the validity of the proposed test. Our motivating application examines the optimal allocation of economic aid in Iraq from February 2007 to July 2008. Drawing on declassified conflict data, we study the weekly assignment of aid projects across districts and find that reallocating aid away from regions with a persistently high level of violence can substantially reduce insurgent attacks.