Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
针对安全表示的优化:激活引导的对抗后缀与拒绝几何
AI总结 研究大语言模型安全表示,引入激活引导的对抗后缀攻击方法,如Activation-Guided GCG和Soft-GCG,发现安全表示分布特点,不同规模模型的抗性不同,结果阐明安全机制编码及破解方式,为设计对齐策略提供指导。
Comments Accepted at the AAAI 2026 Summer Symposium Series. This paper was presented at the ICLR Re-Align workshop under a different name, "Accelerating Adversarial Suffix Optimization via Continuous Relaxation and Activation-Guided Objectives"