arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量大语言模型中的激活控制

Measuring Activation Control in Large Language Models

Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa

arXiv 2608.21664首次发表:更新:

发表机构

UK AI Security Institute(英国人工智能安全研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出Activation Controllability Benchmark量化大语言模型的激活控制能力,发现多数模型可调控残差流,其控制可规避部分激活监控方法,建议追踪该能力以优化模型监控。

AI 中文摘要

安全部署能力日益增强的模型可能会依赖潜在空间监控作为行为评估的补充,尤其是当具备评估感知能力的模型表现出策划或欺骗行为时。然而,如果模型能够控制自身的激活,欺骗行为可能会延伸至潜在空间本身。基于此,我们推出激活可控制性基准(Activation Controllability Benchmark),以量化模型通过自然语言指令调节其残差流(residual stream)的程度。在不同模型家族和能力水平上,我们发现大多数大语言模型(LLMs)能够以一定的时间分辨率控制其残差流激活的方向和幅度,尽管不同模型的表现差异显著。在简单任务中,这种控制水平可以规避基于激活的监控方法(包括线性探针(linear probes)、自然语言自编码器、激活神谕(activation oracles)以及雅可比透镜(Jacobian lens)),尽管规避效果并不完美。这些结果表明,随着内省能力的提升,对激活空间本身的控制可能会成为监控的一个干扰因素;因此,我们建议前沿实验室和评估人员在未来的模型中追踪激活可控制性。

英文摘要

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.

Comments19 figures, 4 tables. Code: https://github.com/mkobalski/activation-control. Data: https://huggingface.co/datasets/joshycodes/activation-control-battery

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑