发表机构
University of Maryland, College Park(马里兰大学学院公园分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出反事实激活潜力(CAP)指标与CSFD算法,揭示LLM中被抑制的安全特征可被越狱利用,补充了激活导向的安全可解释性。
AI 中文摘要
机制可解释性已成为理解大语言模型(LLM)安全行为的主要手段。然而,现有工具主要关注模型的激活神经元或特征。剩余的大量非激活组件的作用对这些方法而言是不可见的。本工作表明,非激活集合包含与拒绝有害提示因果相关的安全关键特征。抑制此类特征可将拒绝转变为顺从,同时逃避主流可解释性工具的检测。我们引入了反事实激活潜力(CAP),该指标将受抑制特征的潜在激活倾向量化为其编码器对齐(输入驱动它的强度)、抑制强度(激活特征抑制它的强度)和安全关键性(拒绝依赖它的程度)的乘积。为了大规模发现被抑制的安全特征,我们提出了CAP引导的安全特征发现(CSFD),一种两阶段过滤算法,无需穷举消融即可从数十万个转码器特征中识别候选安全特征。当候选特征被消融时,相当大比例的试验对有害提示变得顺从。在自然越狱下,作用于最高CAP特征的抑制上升2-4倍,其激活相应下降高达80%。放大特征的抑制器会降低其激活,并提高与受抑制特征相关提示的有害顺从性,而对随机特征则无此效果。我们的实验涵盖五个Gemma、Qwen和Llama模型,参数规模各异。我们的发现表明,越狱可能部分通过抑制安全关键特征而非仅激活有害特征来运作,并且受抑制特征是激活导向的安全行为可解释性的必要补充。
英文摘要
Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.
Comments28 pages, 3 figures, 15 tables. Submitted to ICLR 2027