arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有序结果的正交策略学习

Orthogonal Policy Learning with Ordinal Outcomes

Yue Zhang, Shanshan Luo, Yangbo He

arXiv 2609.22637首次发表:更新:

发表机构

Peking University; Beijing Technology and Business University(北京大学; 北京工商大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对有序结果,提出基于极小极大策略与Neyman正交化的策略学习框架,以处理异质性效用并最小化最坏遗憾,经模拟和SIPP数据验证。

AI 中文摘要

基于条件平均处理效应的策略学习方法在应用于有序结果时,可能会掩盖亚群体异质性。我们开发了一个针对有序结果的策略学习框架,该框架考虑了从治疗中严格获益的个体与未严格获益的个体具有异质性效用的情况。由于在没有进一步条件下,严格获益的概率仅能被部分识别,我们采用极小极大策略,在识别区域上最小化最坏情况下的遗憾值。为了估计由此产生的非光滑目标函数,我们将光滑逼近与Neyman正交化相结合,以消除来自干扰估计的一阶偏差。我们还推导了在受限策略类上的超额最坏情况遗憾界。所提出的方法通过大量模拟实验以及对2022年收入与项目参与调查(SIPP)数据的应用得到了验证。

英文摘要

Policy learning methods based on conditional average treatment effects can obscure subpopulation heterogeneity when applied to ordinal outcomes. We develop a policy learning framework for ordinal outcomes with heterogeneous utilities for individuals who strictly benefit from treatment and those who do not. Since the probability of strict benefit is only partially identified without further conditions, we adopt a minimax strategy which minimizes the worst-case regret over the identification region. To estimate the resulting nonsmooth objective, we combine smooth approximations with Neyman orthogonalization to remove first-order bias from nuisance estimation. We also derive an excess worst-case regret bound over a restricted policy class. The proposed method is validated through extensive simulations and an application to the 2022 Survey of Income and Program Participation (SIPP) data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑