arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

行为的回响:道德历史可以塑造和引导LLM行为选择

Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

Lucio La Cava, Andrea Tagarelli

arXiv 2609.35070首次发表:更新:

发表机构

University of Calabria(卡拉布里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MoralLedger框架,证明LLM的道德历史在行为和表征层面影响决策,且可沿潜在方向干预控制道德选择,为道德审计提供新途径。

AI 中文摘要

对大型语言模型(LLM)的道德评估通常孤立地考虑决策,从而忽略了个人不相关的先前行为是否影响模型后续的选择。这留下了一个问题:道德历史是否以及在多大程度上塑造了LLM的决策行为。先前关于人类道德决策的研究表明,过去的行为可以影响后续的道德选择。基于这一观察,我们以两种互补的方式调查类似效应是否在LLM中出现:在行为层面,通过模型可观察的响应;在表征层面,通过其潜在的内在表征。我们引入了MoralLedger,一个研究行为者的道德历史如何在固定决策情境下塑造LLM行为的框架。在行为层面,我们发现先前的道德历史根据其效价和强度系统地改变后续选择。在内在表征层面,这些历史在残差流中诱导出一个线性可恢复的方向,该方向泛化到保留的示例。沿着这个方向对中性历史提示进行干预,会产生双边的、强度依赖的后续选择变化,其效果强于仅通过提示或通过有利的非道德方向所诱导的效果。据我们所知,这是首次证明行为者先前道德行为的潜在表征可以提供对道德决策的带符号推理时控制。我们的MoralLedger将道德评估扩展到静态困境之外,确立道德历史既是行为敏感性的来源,也是审计和控制LLM道德行为的因果目标。

英文摘要

Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑