arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15354cs.AI

天生不连贯?大型语言模型的道德自我一致性研究

Incoherent by Design? On the Moral Self-Consistency of LLMs

发表机构康奈尔大学 · 康奈尔科技学院 · 南洋理工大学
查看机构详情
  • Cornell University(康奈尔大学)
  • Cornell Tech(康奈尔科技学院)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对GPT、Mistral、Llama等LLM,在义务论等三大伦理框架下,发现其道德推理存在最高78%的矛盾率,内部不连贯是AI对齐的必要前提。

中文摘要 AI 辅助

大型语言模型(LLMs)正越来越多地被用于道德敏感场景,但目前尚不清楚它们是否能在不同情境中一致地应用伦理原则。一个能阐述某项道德原则的模型,在同一情景被重新表述或重构时,仍可能违背该原则。这种不一致性对于任何输出被用于为道德决策提供依据的系统而言都是一个问题。如果生成式系统表现出内部不一致性,那么AI介导系统的认知完整性就会变得不确定。为研究这一问题,我们在受控提示框架内,针对三大主要哲学思想流派——义务论、功利主义和美德伦理学,研究了LLMs道德推理的稳定性。我们构建了多组道德等价情景,其中潜在情境保持不变,仅通过调整表述以反映不同伦理立场和风格扰动。随后我们评估了多个模型的回应,包括GPT、Mistral和Llama。为评估一致性,我们将模型输出转换为结构化逻辑语句,并识别同一哲学流派内生成的回应间的矛盾。我们的结果显示存在显著的不一致性,在不同情景中矛盾率最高达78%。这些发现揭示了生成式AI中更广泛的认知不稳定现象,即模型无法可靠地保持与自身先前输出的一致性。这种不稳定会带来实际后果:由于生成式系统影响人们形成信念、判断行为和吸收价值观的方式,其不一致性也会影响人类的推理和决策。此外,如果一个系统无法一致地表达自身的规范承诺,那么价值对齐就会变成一个移动的目标,而非一个明确定义的目标。因此,我们认为,证明内部不连贯是AI对齐的必要前提。

英文摘要

LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.

补充信息

↑