语言模型中的评估感知:表示、语言化与控制
Evaluation Awareness in Language Models: Representation, Verbalization, and Control
浏览论文内容
中文总结 AI 辅助
该研究系统探测了6个语言模型的评估感知,发现其可线性解码、与语言化部分对齐,且在Olmo模型中随微调放大,需评估时考虑模型内部表示与语言化的脱节。
中文摘要 AI 辅助
能力与安全基准均基于如下假设:接受测试的语言模型的行为可反映其部署时的行为。若模型推断出自身正在接受评估并据此调整响应,则该假设可能不成立。这一被称为“评估感知”的假设,在前沿模型和开放权重语言模型中均被观测到。我们对该现象开展系统研究,在6个语言模型(来自4个家族、3种规模)和3种指标中探测其存在。具体而言,我们研究:(1)处于评估状态是否在模型的激活空间中被线性表示;(2)是否在输出 token 中被语言化(由 LLM-as-judge 评分);(3)是否可通过因果干预调控模型行为。对于开放 checkpoint 的 Olmo 模型,我们还在每个训练阶段测试上述指标。研究发现,所有模型的残差流均可线性解码出评估感知(最佳 AUROC ≥ 0.7)。相比之下,这些表示仅与语言化部分对齐:在部分场景中,二者的相关性和互信息非零,但在模型、层及读取选择间差异显著。不过,沿探测得到的方向调控可改变语言化评分。最后,对 Olmo 各 checkpoint 的比较显示,评估感知在基础模型中已存在,在监督微调各阶段被放大,之后保持稳定;而调控效果则在每个后续训练阶段愈发显著。这些结果表明,评估需考虑模型内部表示、语言化内容及调控间的脱节。
英文摘要
Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and condition their response on such context. This hypothesis, termed ``evaluation awareness'', has been observed in frontier and open-weight language models alike. We provide a systematic study of this phenomenon, by probing for it across six language models (from four families and three sizes) and three metrics. More precisely, we examine whether (i) being under evaluation is linearly represented within the models' activations space, (ii) it is verbalized in their output tokens (as scored by an LLM-as-judge), and (iii) steering causally affects their behavior. For the open-checkpoint Olmo models, we further test these measures at every training stage. In doing so, we report that evaluation awareness is linearly decodable from the residual streams of every model (best AUROC $\geq 0.7$). By contrast, these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices. Nevertheless, steering along probe-derived directions can shift the verbalization scores. Finally, a comparison across the Olmo checkpoints reveals that evaluation awareness is already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter---unlike the effects of steering, that grow more pronounced at every successive training stage. These results show the need for evaluations to account for the disjunction between what models represent internally, what they verbalize, and their steering.