arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于反事实策略的安全贝叶斯优化

Safe Bayesian Optimization with Counterfactual Policies

Katherine Avery, Bruno Castro da Silva, David Jensen

arXiv 2607.05620首次发表:更新:

发表机构

College of Computer Science University of Massachusetts Amherst(计算机科学学院 马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在新干预受安全约束的决策场景下的优化问题,利用共形预测构建反事实基线结果的不确定性区间,集成到安全贝叶斯优化中,确保约束 violation 率可控,还能适应协变量转移,提供了多方面分析。

AI 中文摘要

在许多决策场景中,新干预措施只有在不使结果低于既定阈值时才被接受。例如在临床医学中,新治疗方法只有在不使结果比既定护理标准更差时才常被接受。安全贝叶斯优化在安全约束下最大化目标。在此场景中,安全相对于已知基线策略定义,其结果是反事实的且未被观察到。因此,必须估计基线策略的反事实结果,并使用这些(不确定的)估计来安全地优化目标。我们通过共形预测为反事实基线结果构建有效的不确定性区间来解决估计问题,并展示如何将这些区间集成到安全贝叶斯优化中,以确保违反约束的情况以用户指定的速率或更低发生。我们还展示了如何使这些共形估计适应不同类型的协变量转移。我们提供了安全证明、实验证据和敏感性分析。

英文摘要

In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. For example, in clinical medicine, new treatments are often acceptable only if they do not worsen outcomes relative to an established standard of care. Safe Bayesian optimization maximizes an objective subject to safety constraints. In the setting that we consider here, safety is defined relative to a known baseline policy whose outcomes are counterfactual and therefore unobserved. Thus, the counterfactual outcomes of the baseline policy must be estimated and those (uncertain) estimates must be used to safely optimize the objective. We address this estimation problem by using conformal prediction to construct valid uncertainty intervals for counterfactual baseline outcomes, and we show how these intervals can be integrated into safe Bayesian optimization to ensure that constraint violations occur at or below a user-specified rate. We also show how to adapt these conformal estimates to different kinds of covariate shift. We provide a safety proof, experimental evidence, and a sensitivity analysis.

Comments10 pages main text, 20 pages total

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑