arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24046cs.AI

算法影响揭示对齐的隐藏社会选择结构

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

  • MIT(麻省理工学院)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia

AI总结:

该研究将AI对齐问题转化为凸影响空间上的线性优化,结合福利经济学与机制设计,推导了防策略的社会选择机制及最大化功利主义社会福利的对齐协议,并通过多领域人类偏好实证验证其福利影响。

AI中文摘要:

当AI算法做出影响多个人的决策时,将其对齐成为一个社会选择问题:如何协调并汇总人们关于系统行为的不同偏好,形成一个连贯的单一模型?对齐前沿AI模型的标准方法——基于人类反馈的强化学习——在很大程度上规避了这一问题,且社会选择保障较差。然而,应采用何种替代方案仍不明确。我们表明,通过直接关注算法的福利后果,对齐问题可被重新表述为在凸影响空间上的线性优化,这使其可应用福利经济学和机制设计的标准工具集。该重新表述阐明了对齐协议如何转化为福利后果,反之,社会规划者对福利后果的期望约束如何可转化为对齐协议。我们应用这一转化,证明议题投票和随机独裁机制是防策略且一致的。作为反向应用,我们还利用影响表示推导了一系列对齐协议,这些协议在满足个体或群体伤害限制等各种社会诉求的前提下,最大化功利主义社会福利。我们使用真实人类对肾脏分配、慈善食品分配、大语言模型(LLM)响应及电车难题的偏好,实证说明这些对齐协议的福利影响。

英文摘要:

When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\unicode{x2013}$reinforcement learning from human feedback$\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.

↑