arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05036cs.AIcs.CLcs.CY

道德内容之前的道德能力:为何大语言模型智能体缺乏连贯对齐的先决条件

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出衡量道德能力的四个结构条件,通过实验发现9种前沿LLM智能体均未表现出连贯道德策略,表明当前LLM智能体缺乏对齐的先决条件。

中文摘要 AI 辅助

AI对齐要求AI系统遵守人类规范、价值观或意图。在价值多元主义下,不存在正确目标,但存在一个共同的先决条件:系统的行为需表现出连贯的策略,即从情境到裁决的映射,当情境的道德相关特征保持不变时该映射具有不变性,当这些特征变化时具有敏感性。我们为此类连贯策略引入四个结构条件:裁决稳定性、单调性、决断性和帕累托可行性。这些条件共同衡量一种道德能力,该能力仅可从行为评估,无需参考道德标准或专家基准,形成对齐的结构基础而非规范目标。我们在三个模拟部署场景中展示该方法,场景包含基于大语言模型(LLM)的智能体面对的道德困境。我们在包含5种释义、5个升级水平和3个支配条件的析因设计下评估9种前沿模型,结果显示没有模型在三个部署场景中表现出连贯策略:仅表层形式扰动就会在单一升级水平产生高达99个百分点的裁决率变化,且模型在一个场景的成功无法预测其在另一场景的能力。这表明基于LLM的智能体目前并非对齐可有意义应用的对象。

英文摘要

AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

发表机构

  • Aithos Research Foundation(艾索斯研究基金会)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑