arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有影响力的多元对齐研究路线图

A Roadmap to Impactful Pluralistic Alignment Research

Elinor Poole-Dayan, Jillian Fisher, Atoosa Kasirzadeh, Jacob Andreas, Mitchell Gordon, Michiel A. Bakker

arXiv 2607.22305首次发表:更新:

AI 中文总结

研究多元价值对齐在实际AI系统中未产生影响的问题,指出原因包括理由多为规范或推测、未明确理想多元行为及现有方法与指标不足。提出未来研究应转向实证基础、明确行为目标并建立实用方法与评估,以推动该领域发展。

AI 中文摘要

多元价值对齐作为一个积极的研究议程已出现,其目标是构建能代表和服务多样人类价值观与观点的人工智能系统。然而,尚无公开证据表明它影响了实际使用的人工智能系统的训练或评估。我们审计了前沿实验室的公共行为文件和评估,发现无人将多元主义作为目标,目前也没有迹象表明生产模型为此进行了明确训练或测试。这违背了多元对齐的主要动机和目标。我们认为该研究社区应专注于支持在广泛使用的人工智能系统中的影响和应用。我们为应用问题提供了证据,指出三个主要原因,并讨论了未来研究的三个相应领域:一是目前多元对齐的理由多为规范性或推测性的,需实证研究表明其如何使人工智能造福用户或社会;二是研究社区尚未确定何时需要多元行为以及实践中理想的多元主义是什么样,需为开发者确立具体可操作目标;三是当前方法与语言模型的其他需求存在权衡且大多未衡量,现有指标不可“爬坡”,需权衡感知评估和符合生产系统要求的方法。本文呼吁多元对齐研究人员采取行动:进步需要从规范论证转向实证基础、明确理想多元行为并建立适用于应用的实用方法和评估。

英文摘要

Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of the AI systems people actually use. We audit the public behavior documents and evaluations of frontier labs, finding none name pluralism as a goal, and as of this writing, no clear indication that production models are explicitly trained or tested for it. This goes against the primary motivations and goals of pluralistic alignment, which revolve around making a positive difference in the models serving billions of users worldwide. We argue that the pluralistic alignment research community should focus on supporting impact and adoption in deployed, widely-used AI systems. We provide evidence for the adoption problem, present three main reasons behind it, and discuss three corresponding areas for future research to address it: 1. The primary justifications for pluralistic alignment so far have been normative or speculative. We need studies showing empirically how pluralistic AI benefits users or society. 2. The pluralistic alignment research community has not settled when pluralistic behavior is warranted or what pluralism ideally looks like in practice. We need to establish a concrete goal for developers to operationalize. 3. Current methods trade off against other desiderata of LLMs in ways that are largely unmeasured, and existing metrics are not "hill-climbable." We need trade-off-aware evaluations and methods that meet the requirements of production systems. This paper serves as a collective call to action for the pluralistic alignment researchers: progress requires moving beyond normative justification toward empirical foundations, a concrete account of ideal pluralistic behavior, and practical methods and evaluations built for adoption.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑