发表机构
University of Connecticut; Carnegie Mellon University; Hugging Face(康涅狄格大学; 卡内基梅隆大学; Hugging Face)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出现有LLM道德推理评估聚焦道德价值问题而忽视规范问题,识别出三大缺口并提出含规范表征、标注数据集及分层评估的研究议程,以推动规范推理的系统研究。
AI 中文摘要
近期针对大型语言模型(LLMs)道德能力评估的研究,主要聚焦于我们所称的道德价值问题,即模型输出是否与人类道德价值观相符。相比之下,道德规范问题——也就是模型能否识别并正确应用依赖于具体语境的道德规范——仍未得到充分探索。我们认为这种失衡源于该领域对描述性伦理学框架的依赖,比如道德基础理论和科尔伯格道德发展阶段理论,这些框架更强调价值表征而非规范应用。我们梳理了现有基准与评估方法,发现它们高度集中于价值问题,而关于规范伦理学的讨论则严重不足。我们识别出三个关键缺口:(1)缺乏用于道德规范及其应用的高质量基准数据;(2)对中间推理过程的评估不足;(3)对语境中道德相关特征识别的关注有限。随后,我们提出一项研究议程,包括开发规范理论的标准化形式表征、构建经专家标注的捕捉规范应用的数据集,以及明确区分价值层面与规范层面能力的评估方案。我们的目标是推动对大型语言模型中规范推理的更系统研究。
英文摘要
Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.
Comments8 pages, 1 figure. Accepted for archival publication at the ACL 2026 Workshop on Evaluating Evaluations (EvalEval)