arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28620cs.AIcs.LG

用于策略优化的偏好 elicitation 及其在心脏移植与人类价值观对齐中的应用

Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values

Itai Zilberstein, Ioannis Anagnostides, Zachary W Sollie, Arman Kilic, Tuomas Sandholm

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种用于策略优化的线性效用偏好 elicitation 算法,将其应用于心脏移植分配,通过用户研究学习社区对齐的效用函数,所优化策略的竞争比达0.95,显著优于现状策略。

中文摘要 AI 辅助

偏好 elicitation 对于使 AI 系统与人类价值观对齐至关重要。现有方法(例如用于器官分配的方法)常要求利益相关者比较算法的决策(例如患者 A 与患者 B),这种决策级方法混淆了手段与目的。相反,我们直接针对分配结果 eliciting 偏好以学习用于策略优化的效用函数。我们构建了一种新颖的线性效用偏好 elicitation 算法,在实践中优于现有技术。该算法分为两个阶段:第一阶段通过成对比较学习切割平面,以快速缩小属性权重空间,并通过消除支配区域为第二阶段提供热启动;第二阶段可证明能收敛到用户的效用函数。我们将该技术应用于心脏移植分配,其中策略必须平衡多个竞争目标,如移植后结局、等待名单死亡率、地理便利性和公平性。使用我们的算法,我们开展用户研究以学习和聚合与社区对齐的效用函数,并利用其优化心脏移植策略,使其显著更好地与人类价值观对齐。与事后最优相比,现状策略的竞争比仅为 0.54,而我们的方法接近最优,竞争比为 0.95。

英文摘要

Preference elicitation is essential for aligning AI systems with human values. Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. patient B). Such a decision-level approach conflates the means with the ends. Instead, we elicit preferences directly over allocation outcomes to learn a utility function for policy optimization. We construct a novel preference elicitation algorithm for linear utilities that outperforms prior techniques in practice. Our algorithm has two phases. The first phase learns cutting planes through pairwise comparisons to rapidly shrink the space of possible attribute weights and warm-starts the second phase by eliminating dominated regions. The second phase then provably converges to the user's utility function. We apply our technique to heart transplant allocation where a policy must balance competing objectives such as post-transplant outcomes, waitlist mortality, geographic ease, and equity. Using our algorithm, we conduct a user study to learn and aggregate a community-aligned utility function, and use it to optimize heart transplant policies that are significantly better aligned with human values. Compared to the hindsight optimum, the status quo policy achieves a competitive ratio of just 0.54, while our method is near-optimal with a competitive ratio of 0.95.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Medical University of South Carolina(南卡罗来纳医科大学)
  • Strategy Robot, Inc.(策略机器人公司)
  • Strategic Machine, Inc.(战略机器公司)
  • Optimized Markets, Inc.(优化市场公司)

机构由 AI 辅助整理,请以论文原文为准。

↑