AI 中文总结
本研究通过标注94篇AI价值对齐论文,发现多数未明确定义人类价值,转而依赖偏好,且采用合成数据等自动方法可能关闭价值践行路径,旨在明确该领域哲学承诺以丰富相关辩论。
AI 中文摘要
AI系统能否对齐人类价值?随着大语言模型(LLMs)和多模态基础模型的普及,由有毒言论、幻觉到AI智能体执行未授权行动等危害不断增加。在AI安全领域,这些有害案例常被归为对齐问题,即模型与人类价值不一致。研究者已开展应用和理论层面的AI价值对齐工作,但往往未明确所指的人类价值具体是什么。AI价值对齐领域如何构想人类价值?这些价值构想在技术上如何被操作化和评估?该领域涌现的价值理论对AI的未来意味着什么?我们对94篇价值对齐研究论文进行标注,以辨识其隐含的AI价值理论。多数论文未定义价值,严重依赖偏好作为替代,存在将复杂的文化情境概念简化为二元选择的风险。当研究者放弃使用人工标注者进行模型训练和评估,转而采用合成数据和自动评分器方法来对齐和评估模型时,我们发现这可能会关闭基础模型中用于质疑和践行价值的替代方法。通过明确AI价值对齐的哲学承诺,我们旨在为AI能否以及如何应对人类价值的辩论带来更高的特异性和被忽视的视角。
英文摘要
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.
Comments15 pages, 2 figures