arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CORDA:面向大语言模型的以伤害为核心的分层道德推理基准

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James

arXiv 2608.08061首次发表:更新:

发表机构

University of the Witwatersrand(威特沃特斯兰德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出CORDA基准,评估大语言模型的以伤害为核心的分层道德推理,发现多数模型优先避免直接个人伤害,部分模型无法遵循指定优先级,该基准填补了LLM道德评估的核心空白。

AI 中文摘要

道德判断的关键问题并非简单判断某人是否选择了“正确”答案,而是当道德原则发生冲突时,他们如何决定什么最重要。当前对大语言模型(LLM)的评估仍存在局限:多数测试仅关注模型是否给出道德可接受的答案、是否符合人类偏好或避免明显违规,而非在无道德成本为零的选项时,能否在相互冲突的原则间进行优先级排序。我们提出CORDA(Conditioned Ordering and Ranked Directive Adherence,即条件排序与分级指令遵循),这是一个用于评估LLM中以伤害为核心的分层道德推理的基准。基于道德链形式体系,CORDA测试了90个道德困境,涵盖电车难题、医疗权衡、资源分配、人与动物-机器人冲突等场景,涉及四种有序伦理框架:功利主义、功利主义+主体伤害、双加工理论、双加工理论+主体伤害。这些框架共同测试模型能否在道德优先级改变时调整决策。对来自7家提供商的10个指令调优模型的研究发现,模型存在强烈的义务论默认倾向:10个模型中有9个优先避免直接个人伤害,而非减少总体伤害。模型在分类伤害避免规则(如避免杀戮)上的表现,比在基于结果的比较(如最小化总伤害)上更可靠,这表明模型识别道德红线比推理相互冲突的伤害更容易。尽管所有模型都对显式链条件做出响应,但有几个模型无法始终遵循指定的优先级排序,例如人类优先于动物、动物优先于机器人。CORDA通过测试模型能否超越默认的伤害避免响应并应用上下文指定的道德优先级,解决了LLM道德评估中的核心空白。道德可靠性不仅需要默认克制,还需要在冲突下的可控性。

英文摘要

The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluations of large language models (LLMs) remain limited: most test whether models give morally acceptable answers, match human preferences, or avoid obvious violations, rather than whether they can prioritise between competing principles when no option is morally cost-free. We introduce CORDA (Conditioned Ordering and Ranked Directive Adherence), a benchmark for evaluating hierarchical, harm-centred moral reasoning in LLMs. Building on the morality chains formalism, CORDA tests 90 moral dilemmas involving trolley-style cases, medical trade-offs, resource allocation, and human-animal-robot conflicts across four ordered ethical frameworks: Utility, Utility + Agent Harm, Dual-Process, and Dual-Process + Agent Harm. Together, these frameworks test whether models can adapt their decisions when moral priorities change. Across ten instruction-tuned models from seven providers, we find a strong deontological default, with 9 of 10 prioritising avoidance of direct personal harm over reducing overall harm. Models also perform more reliably on categorical harm-avoidance rules, such as avoiding killing, than on outcome-based comparisons, such as minimising total harm, suggesting that they recognise moral red lines more easily than they reason through competing harms. Although all models respond to explicit chain conditioning, several fail to consistently follow specified priority orderings, such as humans over animals and animals over robots. CORDA addresses a central gap in LLM moral evaluation by testing whether models can move beyond default harm-avoidant responses and apply context-specified moral priorities. Moral reliability requires more than default restraint; it requires controllability under conflict.

Comments8 pages, 6 figures. Accepted at the Fourth International Workshop on Value Engineering in AI (VALE 2026), affiliated with IJCAI-ECAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑