走向多元对齐的路线图
A Roadmap to Pluralistic Alignment
- University of Washington(华盛顿大学)
- Stanford University(斯坦福大学)
- MIT(麻省理工学院)
- Allen Institute for Artificial Intelligence(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出多元对齐路线图,定义三种多元模型与三类基准,论证现有对齐技术限制多元AI,需进一步研究。
AI中文摘要:
随着AI系统能力的增强和普及,设计AI系统以服务于所有人,即具有不同价值观和观点的人群,变得比以往任何时候都更为关键。然而,使模型对齐以服务于多元的人类价值观仍然是一个开放的研究问题。在本文中,我们提出了一条通往多元对齐的路线图,特别以语言模型作为测试平台。我们识别并形式化了在AI系统中定义和实施多元主义的三种可能方式:1) 奥弗顿多元模型,呈现一系列合理响应的光谱;2) 可操控的多元模型,能够转向以反映特定观点;3) 分布多元模型,在分布上对给定人群进行良好校准。我们还形式化并讨论了三种可能的多元基准类别:1) 多目标基准,2) 权衡可操控基准,激励模型转向任意的权衡,以及3) 陪审团多元基准,明确模拟多样化的人类评分。我们利用这一框架论证,当前的对齐技术对于多元AI可能从根本上受限;事实上,我们强调了经验证据,既来自我们自己的实验,也来自其他工作,表明标准的对齐程序可能会减少模型中的分布多元性,从而激发了对多元对齐进一步研究的必要性。
英文摘要:
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.