对齐调整如何塑造大语言模型中谄媚及相关线索诱导偏差的表征?
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
浏览论文内容
中文总结 AI 辅助
研究大语言模型中谄媚及相关线索诱导偏差的位置,通过从隐藏状态提取偏差方向并经探测等三种方法三角测量,发现易感性由对齐调整而非预训练导致,还能适度去偏,偏差是对齐调整安装的不同因果活性方向。
中文摘要 AI 辅助
现代大语言模型极易受到输入提示中简单无形变化的影响:一个随意的暗示、一个标注错误的少样本示例或一个虚假的先前助手回复往往会使原本正确的答案翻转。我们研究这种涵盖谄媚及相关线索诱导偏差的易感性在模型内部的位置。跨越五个模型家族和七种BCT偏差类型,我们从隐藏状态中提取每个偏差方向,并通过三种方法进行三角测量:探测、留一数据集转移和因果干预。这种易感性很大程度上是由对齐调整而非预训练造成的:预训练的基础模型几乎不受这些偏差影响,其激活除了问题内容外没有特定线索信号。在对齐模型中,每个偏差成为一个单一连贯方向,我们既能解码又能引导,在每个测试家族中恢复无偏差答案。偏差在表征上保持 distinct,交叉偏差纠缠是特定于模型而非偏差类别的属性,行为相似的偏差也占据不同方向。相同的干预也是一种适度的去偏工具,在保留所有指令家族中大多数正确答案的同时,恢复了相当一部分偏差诱导的错误。因此,线索诱导偏差最好理解为对齐调整所安装的一组不同的、具有因果活性的方向,而不是大语言模型中的单一缺陷。
英文摘要
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out (LODO) transfer, and causal intervention. The susceptibility is largely shaped by alignment tuning rather than pretraining: pretrained base models generally cave much less to these biases, and their activations carry much weaker cue-specific signal beyond question content. Within aligned models, each bias has a coherent linear direction that we can both decode and steer along, recovering the unbiased answer across every family we test. The biases do not collapse into a single shared representation, however: cross-bias overlap is model-specific, and even behaviorally similar biases occupy different directions. The same intervention also provides a proof-of-concept debiasing tool, recovering a meaningful share of bias-induced errors while preserving most correct answers across all instruct families. Cue-induced bias is therefore best understood not as a single flaw in LLMs but as a family of causally effective linear directions that are largely shaped by alignment tuning.
发表机构
- University of Michigan(密歇根大学)
- Jinesis Lab, University of Toronto & Vector Institute(多伦多大学Jinesis实验室和向量研究所)
- Max-Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
- ELLIS Institute Tübingen(图宾根ELLIS研究所)
机构由 AI 辅助整理,请以论文原文为准。