发表机构
New York University Abu Dhabi; New York University(纽约大学阿布扎比分校; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过分离度量预测语言模型激活引导的可操控性,实验证明其与下游效果强相关,可作为实用诊断工具。
AI 中文摘要
使用一组对比表示来引导语言模型已成为控制模型行为的一种经典且计算高效的方法。尽管在控制某些模型行为方面取得了成功,但激活引导在不同概念上的有效性差异显著;引导向量的泛化特性通常被视为用于构建它们的数据集的函数。我们使这种数据集依赖性主张更加严谨,并表明简单的分离度量在不同设置下与语言模型的下游可操控性强烈相关,即使在控制层和数据集效应之后也是如此。在预测之外,我们从合成叠加实验中提供了证据,表明分离度量与经验特征方向和真实特征方向之间的对齐强烈相关。我们的结果表明,简单的可分离性统计可以作为引导向量可能有效时的实用诊断工具。
英文摘要
Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steerability of language models across diverse settings, even after controlling for layers and dataset effects. Beyond prediction, we provide evidence from a synthetic superposition experiment that separation metrics are strongly correlated with alignment between the empirical and true feature direction. Our results suggest that simple separability statistics can serve as practical diagnostics for when steering vectors are likely to work.