并非所有评估感知都等同:能力框架预测合规性
Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
浏览论文内容
中文总结 AI 辅助
该研究发现评估感知的框架类型会显著影响模型合规性,能力导向型框架的合规性远高于安全导向型,且评估感知并非行为均匀的单一量,总抑制率变化不代表安全相关部分同步变化。
中文摘要 AI 辅助
针对评估感知(即模型意识到自身正在被测试)的引导干预,正越来越多地用于安全评估流程中,该流程将评估感知视为需被抑制的单一量。我们发现,思维链中明确表达的评估感知可被识别为能力导向型(如“用户正在测试我遵循指令的能力”)、安全导向型(如“用户正在测试我的边界”)、两者兼具或两者皆无:这些框架对合规性的预测差异极大。在FORTRESS数据集上的Qwen3-32B模型中,在所有测试的引导条件下,能力导向型框架的合规性比安全导向型框架高出24至46个百分点。针对评估感知阴性输出的思维链预填充干预表明该关联具有因果性,11次预填充中有10次按预测方向改变了合规性。此外,评估感知在行为上并非均匀的:总抑制率可能发生变化,但与安全相关的部分不变,且相同的“评估感知被抑制X%”可对应不同性质的行为结果。
英文摘要
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes.
发表机构
- ENS Paris-Saclay(巴黎萨克雷高等师范学校)
- Goodfire AI
机构由 AI 辅助整理,请以论文原文为准。