AI 中文总结
本研究通过机制可解释性分析,在LLaMA-2-7B-chat-hf中识别出负责越狱响应的关键计算电路,剪枝这些电路可降低80%攻击成功率,为对齐LLMs的防御提供了可解释的新方向。
AI 中文摘要
尽管经过广泛的安全对齐,大语言模型(LLMs)仍易受越狱攻击,这类攻击会绕过安全防护以诱导生成有害内容。此前研究将该漏洞归因于安全训练的局限性,但LLMs处理对抗性提示的内部机制仍知之甚少。我们针对一款大规模安全对齐LLM(聚焦LLaMA-2-7B-chat-hf)的越狱行为开展机制分析,借助边缘归因修补和子网络探测,系统识别出对越狱提示生成肯定性响应的计算电路。在首个token预测阶段剪枝这些电路,可将攻击成功率降低最多80%,证明其在安全绕过中的关键作用。我们的分析揭示了介导对抗性提示利用的关键注意力头和MLP通路,阐明重要token如何通过这些组件传播以覆盖安全约束。这些发现加深了对对齐LLMs中对抗性漏洞的理解,并为基于机制可解释性的针对性、可解释防御机制铺平了道路。
英文摘要
Despite extensive safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts remain poorly understood. We present a mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf. Leveraging edge attribution patching and subnetwork probing, we systematically identify computational circuits responsible for generating affirmative responses to jailbreak prompts. Ablating these circuits during the first token prediction can reduce attack success rates by up to 80\%, demonstrating its critical role in safety bypass. Our analysis uncovers key attention heads and MLP pathways that mediate adversarial prompt exploitation, revealing how important tokens propagate through these components to override safety constraints. These findings advance the understanding of adversarial vulnerabilities in aligned LLMs and pave the way for targeted, interpretable defense mechanisms based on mechanistic interpretability.
CommentsAccepted at the 2nd Workshop on Reliable and Responsible Foundation Models of the 42nd International Conference on Machine Learning (ICML) 2025