arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预测激活引导的副作用

Forecasting Side Effects of Activation Steering

Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun

arXiv 2608.11227首次发表:更新:

AI 中文总结

该研究针对激活引导易产生意外副作用的问题,构建交叉效应矩阵揭示副作用的系统性与可预测性,为激活引导的安全部署提供支持。

AI 中文摘要

激活引导通过向语言模型的隐藏激活添加学习到的方向来修改模型,无需重新训练即可实现针对性的行为改变。尽管有效,但引导往往会对其他行为产生意外副作用,导致难以安全部署。因此我们提出问题:能否在应用引导前预测这些副作用?我们通过在三个开源权重语言模型的67种行为分类上构建交叉效应矩阵来回答该问题。研究发现副作用普遍存在、具有结构性且常呈不对称性,揭示了现有基于相似性的启发式方法无法解释的交互作用。尽管存在这种复杂性,我们表明副作用在执行引导前大多可预测:其幅度主要取决于目标行为,而其方向可从模型未引导的表示中预测,准确度远高于简单基线。我们的结果证明激活引导具有系统性且可预测的副作用,支持主动安全审计并推动引导干预的更明智部署。

英文摘要

Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model's unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.

Comments24 pages, 5 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑