发表机构
Institute of Astronomy & Kavli Institute for Cosmology, University of Cambridge; Risk and Security AI Lab, Visa Inc.(剑桥大学天文学研究所及卡夫利宇宙学研究所; Visa公司风险与安全人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文分析激活引导的去偏方向,发现其主要由模型置信度主导而非公平性,引导降低偏差是通过降低置信度甚至导致弃权,提示基于引导的去偏结果需谨慎解读。
AI 中文摘要
激活引导(activation steering)作为一种轻量级的推理时去偏技术,在大语言模型中越来越受欢迎。然而,先前的研究报告指出,引导向量泛化能力差,对模型性能产生意外影响,并且在新数据集上的迁移有限。我们的工作分析了用于激活引导的去偏方向实际编码了什么,以揭示其性能不一致的原因。我们研究了通过对比反偏提示和偏置提示的激活而获得的线性去偏方向,并将其作为引导干预在偏差和通用知识基准上进行了评估。我们发现,该方向主要由模型置信度主导,在激活空间中从高概率令牌区域指向低概率令牌区域,而非编码模型偏差的有意义表示。沿该方向引导确实降低了测量的偏差,但这是降低模型置信度的结果:在问答基准上,我们发现这种引导导致模型弃权(不执行),副作用是提高了公平性指标。我们的实验表明,模型置信度是隐藏空间中偏置提示和反偏提示之间的主要分离因素,表明隔离与模型置信度解耦的线性偏差表示是困难的,基于引导的去偏结果应谨慎解读。简而言之,引导似乎减少了偏差,不是通过纠正模型的潜在偏好,而是通过降低其置信度,即使在与偏差无关的任务上也是如此。
英文摘要
Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model's underlying preferences, but by making it less confident, even on tasks unrelated to bias.