AI 中文总结
研究大语言模型安全表示,引入激活引导的对抗后缀攻击方法,如Activation-Guided GCG和Soft-GCG,发现安全表示分布特点,不同规模模型的抗性不同,结果阐明安全机制编码及破解方式,为设计对齐策略提供指导。
AI 中文摘要
大语言模型中的行为对齐常常掩盖脆弱的内部安全表示。近期工作表明拒绝行为由激活空间中的低维方向介导。这引发了关于此类表示如何通过优化进行构建、定位和访问的问题。我们研究对抗后缀攻击作为表示对齐的一种探测方式。我们引入激活引导的广义对抗梯度(Activation-Guided GCG),它用直接针对模型内部拒绝方向的损失取代基于输出的目标。在多个目标变体中,我们发现全局抑制所有层和位置的拒绝比针对单个层-位置对更有效。这表明安全表示分布在前向传播中而非因果性地局限于单个位点。我们还引入软广义对抗梯度(Soft-GCG),它是使用Gumbel-Softmax对离散后缀优化的连续松弛。Soft-GCG在提高攻击成功率的同时比标准GCG加速了33倍。在不同模型规模上评估发现,在我们计算受限的设置下,较小模型仍然易受攻击,而较大模型能抵抗基于激活和后缀的攻击,这与经过更好安全训练的更大模型更难被破解一致。我们的结果阐明了当代模型中安全机制如何编码以及如何被破解。这些见解为设计更强大且具有表示感知的对齐策略提供了具体指导。
英文摘要
Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured, localized, and accessed by optimization. We study adversarial suffix attacks as a probe of representational alignment. We introduce Activation-Guided GCG, which replaces output-based objectives with losses that directly target a model's internal refusal direction. Across several objective variants, we find that suppressing refusal globally across all layers and positions is more effective than targeting a single layer-position pair. This suggests that safety representations are distributed across the forward pass rather than causally localized to a single site. We further introduce Soft-GCG, a continuous relaxation of discrete suffix optimization using Gumbel-Softmax. Soft-GCG achieves a 33 $\times$ speedup over standard GCG while improving attack success rates. Evaluating across model scales, we find that smaller models remain vulnerable while larger models resist both activation- and suffix-based attacks at our compute-constrained settings, consistent with larger and better safety trained models being harder to jailbreak. Together, our results clarify how safety mechanisms are encoded and can be broken in contemporary models. These insights provide concrete guidance for designing more robust and representation-aware alignment strategies.
CommentsAccepted at the AAAI 2026 Summer Symposium Series. This paper was presented at the ICLR Re-Align workshop under a different name, "Accelerating Adversarial Suffix Optimization via Continuous Relaxation and Activation-Guided Objectives"