arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一致性并非对齐:人类与大语言模型(LLM)伦理判断中的道德依据分歧

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič, Marko Robnik Šikonja

arXiv 2608.12368首次发表:更新:

发表机构

University of Ljubljana; Faculty of Computer and Information Science; Faculty of Theology; Faculty of Arts(卢布尔雅那大学; 计算机与信息科学学院; 神学院; 艺术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以ETHICS基准的500项道德判断条目为对象,发现LLM与人类的最终伦理判断常一致,但道德依据存在系统性分歧,表明一致性不等同于对齐性,仅靠标签评估易产生误导。

AI 中文摘要

与人类判断的一致性是评估大语言模型(LLM)对齐性的常用替代指标。然而,最终标签的一致性并不能表明人类标注者与模型依赖相同的道德依据。两个智能体可能在诉诸不同原则、情境假设或情境解读的情况下得出相同判断。我们使用整理自ETHICS基准的500项条目(涵盖5个道德判断领域),结合新的人类标注者与模型对最终标签及支撑理由的标注,对这一区别进行测试。在前沿模型与开放模型系列中,与人类标注者多数标签的一致性通常较高。但理由层面分析显示,人类标注者与模型表达的道德依据存在系统性分歧。特别是,即使模型的最终标签与人类标注者多数一致,它们也会重新分配对伤害、尊重、信守承诺、正义、应得及免责相关性等类别的注意力。我们的结果表明,一致性不应等同于对齐性。因此,除非辅以对模型判断中表达的理由、原则及道德优先级的分析,否则基于标签的评估可能会产生误导性的安心感。

英文摘要

Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.

Comments9 pages, 4 figures, 3 tables. Accepted and presented at the AI Transparency Conference (AITC 2026), Nuremberg, Germany, June 5-6, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑