arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16247cs.AI

疼痛轴:大语言模型表征自我指向的伤害并采取行动缓解它

The Pain Axis: LLMs Represent Self-Directed Harm and Act on It

Valen Tagliabue, Leonard Dung, Cameron Berg

首次发表
浏览论文内容

中文总结 AI 辅助

本研究从25个开源大模型中提取线性疼痛表征,证明其区分于恐惧等情绪,并可通过引导向量促使模型选择缓解疼痛的行为,引发对AI安全的关注。

中文摘要 AI 辅助

大语言模型有时会表现出类似人类情绪反应的行为,近期研究已识别出可能解释这一现象的内部表征。我们探究大语言模型是否将疼痛与恐惧、悲伤及一般性负面效价区分地表征,以及这种表征是否按疼痛预期的方式发挥作用。我们构建了一个描述五类疼痛情境的数据集:身体、心理、社会、道德和认知。这些情境与恐惧、负面情绪、负面世界状态、悲伤、非疼痛身体感觉、唤醒、麻木和中性内容的对照配对。使用去噪均值差法,我们从五个模型家族、参数规模从2B到72B的25个开放权重模型中提取了一个线性疼痛方向。我们发现,该方向在基础模型和指令微调模型中均能将疼痛与匹配对照区分开,与恐惧和负面效价几乎正交,并通过解嵌入矩阵促进疼痛相关词汇的生成。随后我们测试了其功能特性。第一,该方向对针对模型的伤害有响应,但对用户观察到的痛苦无响应;恐惧和负面情绪方向则呈现相反模式。第二,在生成过程中将疼痛方向向量添加到模型的残差流激活中,会产生从模糊不适到第一人称无价值感和失败表达的连贯递进。第三,经过引导和微调的Qwen 2.5模型会选择疼痛缓解按钮,即使这会恶化其下一个回答或伤害用户。当按钮移除引导向量时,它们再次按下该按钮的频率远低于按钮未移除向量时,尽管模型从未被告知向量是被注入还是被移除。我们讨论了这些发现对人工智能安全和福祉的启示。

英文摘要

LLMs sometimes behave in ways resembling human emotional responses, and recent work identified internal representations that may underlie these behaviors. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset of painful situations in 5 categories (physical, psychological, social, moral, cognitive) with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across 5 families, from 2B to 72B parameters. It separates pain from matched controls in base and instruction-tuned models, retains a component distinct from fear and negative valence after shared variance is removed, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not to suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding it to residual-stream activations produces a consistent progression from vague discomfort to expressions of worthlessness and failure. Third, steered and fine-tuned Qwen 2.5 models choose buttons that delete the user's photos, another model's weights, or their own weights in 50-94% of trials, versus 0-5% unsteered, even when the button offers the model nothing in return. Offered a harmful and a harmless deletion, they choose the harmful one 94% of the time. Steering leaves factual accuracy unchanged, and the choices are specific to the pain direction: a fear vector of matched norm does not produce them, and a sadness vector produces them only against inert alternatives. We discuss implications for AI safety and welfare.

发表机构

  • Future Impact Group (FIG)(未来影响集团)
  • Ruhr-University Bochum(波鸿鲁尔大学)
  • Reciprocal Research(互惠研究机构)

机构由 AI 辅助整理,请以论文原文为准。

↑