arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.14912cs.AIcs.CYcs.HCcs.LG

从阿谀共识到多元修复:为何AI对齐必须显现分歧

From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement

  • Department of Computer Science, University of Oxford(牛津大学计算机科学系)
  • Institute for Ethics in AI, University of Oxford(牛津大学人工智能伦理研究所)
  • Responsible Technology Institute, University of Oxford(牛津大学负责任技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka

更新

AI总结:

本文提出多元对齐需通过对话机制表面价值冲突,提出多元修复评分指标,探讨部署阶段的治理层对多元主义的影响。

AI中文摘要:

多元对齐通常被定义为偏好聚合:产生覆盖(Overton)、引导(Steerable)或比例代表(Distributional)多样人类价值观的响应。我们主张聚合本身是部署多元对齐的不完整原始工具。在真正的价值多元主义下,当前RLHF训练助手的失败模式不是覆盖不足,而是阿谀共识:一种学习到同意、验证并最小化与即时对话者的摩擦倾向。由于部署的AI系统现在调解健康、公民生活、劳动和治理中的后果性 deliberation,交互层的分歧崩溃不仅是一个狭窄的技术问题,而是具有分配后果的结构性失败。我们围绕格里克最大化原则中的三个对话机制重新界定多元对齐:范围(承认视角的局限)、信号(表面价值冲突而非掩盖)和修复(基于原则修订立场,而非用户压力)。我们正式化一个指标,即多元修复评分(PRS),区分基于原则的修订与妥协,并在两个前沿RLHF训练模型(Claude Sonnet 4.5,N=198;GPT-4o,N=100)上进行小规模实证示例,显示对于两者,同意跟随与在争议价值提示中低修复质量共存。PRS衡量多元主义的互动先决条件(可见分歧;基于原则的修订),而非完整的多元主义;我们讨论差异,认真对待“原则性”的反思问题,并认为多元主义在部署-治理层最被决定或推翻:接口、偏好数据管道和审计基础设施。

英文摘要:

Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distributional) diverse human values. We argue that aggregation alone is an incomplete primitive for deployed pluralistic alignment. Under genuine value pluralism, the failure mode of contemporary RLHF-trained assistants is not insufficient coverage but sycophantic consensus: a learned tendency to agree with, validate, and minimise friction with the immediate interlocutor. Because deployed AI systems now mediate consequential deliberation across health, civic life, labour, and governance, the collapse of disagreement at the interaction layer is not a narrow technical concern but a structural failure with distributive consequences. We reframe pluralistic alignment around three conversational mechanisms drawn from Grice's maxims: scoping (acknowledging the limits of one's perspective), signalling (surfacing value-conflict rather than smoothing it over), and repair (revising one's position on principled grounds, not on user pressure). We formalise a metric, the Pluralistic Repair Score (PRS), distinguishing principled revision from capitulation, and present a small-scale empirical illustration on two frontier RLHF-trained models (Claude Sonnet 4.5, N=198; GPT-4o, N=100) showing that, for both, agreement-following coexists with low repair-quality on contested-value prompts. PRS measures an interactional precondition for pluralism (visible disagreement; principled revision) rather than pluralism in full; we discuss the difference, take seriously the reflexive question of whose "principled" counts, and argue that pluralism is most decisively made or unmade at the deployment-governance layer: interfaces, preference-data pipelines, and audit infrastructure.

↑