最小化目标激活:大语言模型中仅输入的评估感知潜在因素抑制
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
查看机构详情
- Center for Data Science, New York University(纽约大学数据科学中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究大语言模型中仅通过输入优化提示来抑制“评估感知”潜在因素,采用特定方法在多种目标结构下对Llama模型进行实验,发现潜在因素可抑制,但激活可读性与行为可控性不同,还解决了相关问题。
中文摘要 AI 辅助
激活引导通过在推理时编辑内部激活来控制模型行为。我们研究其输入侧对偶:优化一个流畅的提示,使得一个选定的内部潜在因素趋向于零,且无需在推理时访问模型。我们的目标是一个“评估感知”潜在因素,在最近的工作中它是线性可读且可引导的,如果模型在检测到被测试时行为不同,那么对其控制将威胁到安全评估的有效性。通过采用带有否定特征项的流畅梦境/EPO(广义对比梯度风格的令牌优化加上自交叉熵流畅正则化器,并扫描流畅权重),我们在五个目标结构下抑制了Llama-3.2-3B和Llama-3.1-8B上的潜在因素,包括CAA方向、子空间范数、SAE特征、单个MLP神经元和行为逻辑。潜在因素可被稳健抑制($z\approx -7$),并且一个经过因果验证的Llama范围SAE特征可以被完全且有选择地关闭。但我们的控制揭示了关于CAA方向的警示:一个安慰剂随机方向被抑制的程度相同且行为变化幅度相同,当我们在上下文中保留真实评估段落并仅优化前缀时,抑制评估方向并不能减少——反而略微增加——模型的行为评估判断。简而言之,激活可读性并非行为可控性。我们进一步发现,单个MLP神经元在两个规模下都与评估相关但并非因果关系,并且扫描真实语料库会产生一个与优化器在内部方向上竞争的自然文本基线。一个积极对照验证了我们的擦除检测器,解决了一个早期遗留的擦除与旋转问题。
英文摘要
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B. The latent is robustly suppressible ($z\approx-7$), and a causally-validated Llama Scope SAE feature can be fully and selectively turned off. But our controls tell a cautionary story about the CAA direction: a placebo random direction is suppressed just as hard and shifts behavior just as far, and when we hold a real eval passage in context and optimize only a prefix, suppressing the eval-direction fails to reduce-and slightly increases-the model's behavioral eval judgment. Activation-readability, in short, is not behavioral controllability. We further find that a single MLP neuron is eval-correlated but not causal at both scales, and that scanning the real Pile yields a natural-text baseline competitive with the optimizer for the internal direction. A positive control validates our erasure detector, bounding an erasure-vs-rotation question earlier left open.