arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05578cs.CR

通过激活分析检测语言模型中的安全训练修改

Detecting Safety Training Modification in Language Models via Activation Analysis

Glen Messenger

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出AMS工具,通过激活空间几何结构检测语言模型的安全训练修改,在14种模型配置上验证了其有效性,可区分四类安全训练修改中的前三者。

中文摘要 AI 辅助

我们提出了AMS(基于激活的模型扫描器),这是一种通过测量激活空间中安全相关概念的几何结构来检测语言模型安全训练修改的工具。安全训练会在有害内容和良性内容类别之间形成可测量的区分;某些安全修改会使该结构崩溃或旋转,而其他修改则会保持其完整性。我们在涵盖4个架构家族(Llama、Gemma、Qwen、Mistral)的14种模型配置以及4类安全修改类别(指令微调、基础模型、abliterated模型、未审查微调)上验证了AMS。阈值的留一法交叉验证准确率达到71%(10/14);sigma点估计的自助法95%置信区间的中位宽度为3.4 sigma。我们针对每个模型在20个分层JailbreakBench提示上测量行为合规性,发现有害内容概念的sigma与合规性的Pearson相关系数r=-0.546(p=0.043),呈方向性但存在显著噪声。机制分析确定了由激活空间特征区分的四类安全训练修改:训练移除会使聚类分离崩溃;权重正交化abliteration既会使分离崩溃又会旋转拒绝方向;无崩溃的旋转abliteration在保持分离的同时旋转方向;行为微调则同时保持幅度和方向。AMS的一级sigma阈值可检测前两类;二级方向相似性验证可检测第三类;第四类无法通过仅激活探测检测,代表一种已记录的失败模式。我们讨论了阈值校准、单次运行测量的局限性以及仅行为安全修改检测的开放问题。

英文摘要

We introduce AMS (Activation-based Model Scanner), a tool that detects modifications to safety training in language models by measuring the geometric structure of safety-relevant concepts in activation space. Safety training creates measurable separation between harmful and benign content classes; certain safety modifications collapse or rotate this structure, while others leave it intact. We validate AMS across 14 model configurations spanning 4 architecture families (Llama, Gemma, Qwen, Mistral) and four safety-modification categories (instruction-tuned, base, abliterated, uncensored fine-tunes). Leave-one-out cross-validation of thresholds achieves 71% accuracy (10/14); bootstrap 95% confidence intervals on sigma point estimates have median width 3.4 sigma. We measure behavioral compliance on 20 stratified JailbreakBench prompts per model and find that sigma on the harmful-content concept predicts compliance with Pearson r = -0.546 (p = 0.043), directionally but with meaningful noise. Mechanistic analysis identifies a four-class taxonomy of safety-training modifications distinguished by activation-space signature: training removal collapses cluster separation; weight-orthogonalization abliteration both collapses separation and rotates the refusal direction; rotation-without-collapse abliteration preserves separation while rotating direction; and behavioral fine-tuning preserves both magnitude and direction. AMS's Tier 1 sigma-threshold detects the first two classes; Tier 2 direction-similarity verification detects the third. The fourth is undetectable by activation-only probing and represents a documented failure mode. We discuss threshold calibration, limitations of single-run measurement, and the open problem of detecting behavioral-only safety modifications.

补充信息

↑