稳定策略学习
Stable Policy Learning
浏览论文内容
中文总结 AI 辅助
本文研究策略学习如何平衡预期福利与抽样风险,提出策略投票袋装法,通过跨子样本平均投票保持福利并降低风险,并给出严格理论界限。
中文摘要 AI 辅助
在基于证据的政策制定中,通常观察一个实验样本,然后大规模实施学到的策略建议。从实验数据中学习到的策略在预期福利方面可能表现良好,但实验中的随机抽样可能产生福利结果较差的建议。在本文中,我们提出一个问题:策略学习算法应如何平衡预期福利与抽样风险?我们的主要贡献是表明算法稳定性在表征和应对这一权衡中起着核心作用。直观地说,如果策略学习算法的建议在一个实验单元被替换时保持稳定,那么该算法的抽样风险有限。我们提出了一种称为策略投票袋装法的策略学习方法,该方法在许多子样本上学习治疗决策,然后将它们的投票平均为治疗概率。相对于使用一个子样本,跨子样本平均可保持预期福利,并提高风险厌恶研究者的预期效用。我们推导了将估计精度、子样本大小和福利变化联系起来的严格界限,包括在CARA效用下的精确保证。
英文摘要
In evidence-based policymaking, typically one experimental sample is observed, then a learned policy recommendation is implemented at scale. Policies learned from the experimental data can perform well in expected welfare, yet random sampling in the experiment can produce recommendations with poor welfare outcomes. In this paper, we ask: how should policy learning algorithms balance expected welfare against sampling risk? Our main contribution is to show that algorithmic stability plays a central role in characterizing and navigating the tradeoff. Intuitively, if a policy learning algorithm's recommendation remains stable when one experimental unit is replaced, then that algorithm has limited sampling risk. We propose a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities. Relative to using one subsample, averaging across subsamples preserves expected welfare and improves expected utility for a risk-averse researcher. We derive sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.
发表机构
- Harvard University(哈佛大学)
- University of Mannheim(曼海姆大学)
机构由 AI 辅助整理,请以论文原文为准。