arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20505stat.ME

政策干预的反事实优化:词汇排序与跨越式发展

Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging

Martina Scauda, Tobias Freidling, Qingyuan Zhao

AI总结:

该研究针对数据驱动政策学习忽略个体伤害的问题,提出反事实政策优化方法,确定最优政策转换的词汇跨越式结构,在I-SPY2试验中验证其可改变治疗-亚组优先级结论。

AI中文摘要:

大多数数据驱动的政策学习方法会最大化平均结果,却忽略了平均有益的政策仍可能对相当一部分个体造成伤害的可能性。受“不造成伤害”这一伦理原则的驱动,我们研究如何设计从基线政策出发的改变,以提升整体福利,同时将个体伤害的最坏情况概率或期望控制在指定限值以下。我们确定了最优政策转换具有词汇跨越式结构的充分条件:由协变量和当前治疗定义的群体按优先级得分排序,任何治疗变更都会将其直接转移到条件最优治疗。我们在潜在结果间依赖关系的多种模型下推导了该得分,并通过对I-SPY2乳腺癌平台试验的重新分析,展示了这种感知伤害的政策优化方法,表明考虑反事实伤害可能会导致关于哪些治疗-亚组对在进一步临床评估中应被降级的不同结论。

英文摘要:

Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of "first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.

↑