AI 中文总结
该研究发现,道德AI偏好征集流程中的特征范围、投票者构成、问题措辞三项关键选择会显著影响AI决策,仅靠投票汇总无法实现公平透明的AI,需对流程各阶段审计披露。
AI 中文摘要
随着人工智能系统在社会中做出更多带有道德负荷的决策,一种应对方式是道德偏好 elicitation( elicitation 译为“ elicitation 即偏好 elicitation,此处保留术语,首次出现可理解为偏好征集)。在该方法中,研究人员就假设困境对参与者进行民意调查,并使用汇总的投票结果训练策略,随后由人工智能模型大规模应用该策略。在任何投票开始前,开发者会在道德人工智能偏好征集流程中做出三项关键选择:特征范围界定、投票者抽样和问题框架构建。换言之,他们决定哪些特征需提交投票、纳入哪些投票者以及如何呈现问题。这些选择往往不透明、未被记录,且被视为技术细节而非规范性问题。我们在一项常见的实证研究中考察了每一项选择,证明每一项选择都能塑造道德人工智能偏好征集产生的偏好。在三个部署场景(即人工智能肾脏分配、模拟缺勤工人的 AI agents(AI agents 译为“智能体”)以及描绘逝者的生成式人工智能)中,分两个阶段开展研究(N = 809),我们考察了道德人工智能偏好征集流程的三个主要阶段。首先,与道德相关的特征会随场景变化,这表明不应假设特征框架可跨部署领域迁移。其次,约三分之一特征的偏好会因政治意识形态而异,部分差异方向相反,因此投票者群体的意识形态构成会影响最终的汇总偏好概况。第三,偏好征集问题的措辞可将意识形态差距缩小或扩大达整整一个量表单位,框架条件还会改变道德基础与参与者判断的关联方式。综上,这些发现表明,基于投票的对齐仅靠汇总无法实现公平或透明的人工智能;至少,应对道德人工智能偏好征集流程的每个阶段进行审计并公开披露。
英文摘要
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.
Comments35 pages, 11 Figures