多轮攻击中危害性的几何结构
The Geometry of Harmfulness in Multi-Turn Attacks
浏览论文内容
中文总结 AI 辅助
本研究通过分析多轮攻击中LLM隐藏状态的几何动态,发现危害性表示随轮次增加而线性可分且与拒绝表示弱对齐,解释了单轮防御失效的原因,并提出防御需考虑时间动态。
中文摘要 AI 辅助
大型语言模型(LLMs)仍然容易受到绕过安全对齐以诱导有害输出的对抗性攻击。目前尚不清楚在多轮攻击过程中,危害性和拒绝表示如何演变,以及为什么单轮防御在多轮设置中效果较差。本研究探讨了危害性和拒绝表示的几何结构及时间动态如何跨多轮攻击演变。我们使用三种多轮攻击框架(Crescendo、ActorAttack和X-Teaming)分析了三个指令微调LLM(Llama-3.1-8B-Instruct、Qwen2.5-7B-Instruct和Gemma-2-9B-it)的隐藏状态表示,并检查了在不同上下文配置下,表示在对话轮次、模型层和令牌位置上的行为。跨模型和框架,我们发现:(1)每个攻击框架遍历不同的几何方向,但每个框架在诱导有害输出方面都取得了相当的成功;(2)多轮危害性方向在中间到后期模型层的轮次结束令牌位置上,随着轮次增加变得越来越线性可分;(3)危害性表示与拒绝相关表示的弱对齐。结果表明,多轮攻击并非通过抑制模型内部的危害性表示来成功。相反,危害性表示在对话轮次中变得越来越可分,同时与拒绝相关表示仅保持弱对齐。这些发现是静态单轮安全探针在多轮设置中可能性能下降的一种可能解释,并表明稳健的防御必须考虑时间表示动态,而不是用孤立或单轮提示来识别危害性。
英文摘要
Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work investigates how the geometry and temporal dynamics of harmfulness and refusal representations evolve across multi-turn attacks. We analyzed hidden-state representations from three instruction-tuned LLMs (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-it) using three multi-turn attack frameworks (Crescendo, ActorAttack, and X-Teaming), and examined representation behavior across conversation turns, model layers, and token positions under various context configurations. Across models and frameworks, we found that (1) each attack framework traverses different geometric directions, yet each achieves comparable success in eliciting harmful outputs; (2) multi-turn harmfulness directions became increasingly linearly separable at the end-of-turn token position across turns in middle to late model layers; and (3) harmfulness representations are weakly aligned with refusal-related representations. The results indicate that multi-turn attacks do not succeed by suppressing the model's internal representation of harmfulness. Instead, harmfulness representations become increasingly separable across conversation turns, while remaining only weakly aligned with refusal-related representations. The findings are one possible explanation for why static single-turn safety probes may degrade in multi-turn settings, and suggest that robust defenses must consider temporal representation dynamics rather than identifying harmfulness with isolated or single-turn prompts.
发表机构
- Applied AI Research(应用人工智能研究院)
- TELUS Digital(研科数字)
机构由 AI 辅助整理,请以论文原文为准。