从检测到弃权:通过电路引导的权重缩放实现更安全的大语言模型(LLMs)
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
查看机构详情
- University of California San Diego(加利福尼亚大学圣迭戈分校)
- Halıcıoğlu Data Science Institute, University of California San Diego(加利福尼亚大学圣迭戈分校哈利乔格鲁数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文从机制可解释性视角刻画LLM的多阶段安全电路,通过电路引导的权重缩放提升6个LLM在对抗攻击下的安全率26.5%,仅造成1.7%的准确率下降,为LLM安全提供了电路级解释。
中文摘要 AI 辅助
尽管已开展大量对齐工作,大语言模型(LLMs)在对抗性提示下仍易生成不安全内容,但其安全行为的内部机制却鲜为人知。本文从机制可解释性视角研究LLM安全,刻画了组织弃权(不执行)行为的多阶段安全电路,包含:(i)对有害输入作出响应的有害检测头(Harmful Detection Heads);(ii)在残差流中调解并稳定安全信号的安全神经元(Safety Neurons);(iii)将这些信号转化为安全响应生成的弃权头(Refusal Heads)。通过针对性的注意力头和神经元级干预,本文提供了与该电路组织一致的因果证据,表明抑制上游有害检测头会破坏下游弃权行为,且安全神经元介导了这种交互。本文验证了该分解模式在多种LLM架构和对抗性攻击场景中均存在,并采用简单的、架构保留的权重缩放作为机制探针测试其功能相关性。在6个LLM上,电路引导的缩放使攻击下的安全率提升26.5%,同时在4个标准基准上仅产生1.7%的准确率下降。总体而言,本文的结果支持LLM安全的电路级解释,并表明机制抽象可揭示对齐行为背后稳定且可迁移的模式。
英文摘要
Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage *safety circuit* that organizes refusal behavior, consisting of (i) $\textbf{Harmful Detection Heads}$ that respond to harmful inputs, (ii) $\textbf{Safety Neurons}$ that mediate and stabilize safety signals in the residual stream, and (iii) $\textbf{Refusal Heads}$ that translate these signals into safe response generation. Using targeted attention-head and neuron-level interventions, we provide causal evidence consistent with this circuit organization, showing that suppressing upstream Harmful Detection Heads disrupts downstream refusal behavior and that safety neurons mediate this interaction. We validate that this decomposition recurs across multiple LLM architectures and adversarial attack settings, and use simple, architecture-preserving weight scaling as a mechanistic probe to test its functional relevance. Across six LLMs, circuit-guided scaling improves safety rates under attacks by 26.5%, while incurring only a 1.7% accuracy drop across four standard benchmarks. Overall, our results support a circuit-level interpretation of LLM safety and suggest that mechanistic abstractions can reveal stable and transferable patterns underlying aligned behavior.