arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

AAAI Conference on Artificial Intelligence · 会议 · Artificial Intelligence

2026-07-13 至 2026-07-13 共收录 1
2607.08883 2026-07-13 cs.LG 新提交

Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

针对安全表示的优化:激活引导的对抗后缀与拒绝几何

Ege Çakar, Hannah Guan, Kayden Kehe

AI总结 研究大语言模型安全表示,引入激活引导的对抗后缀攻击方法,如Activation-Guided GCG和Soft-GCG,发现安全表示分布特点,不同规模模型的抗性不同,结果阐明安全机制编码及破解方式,为设计对齐策略提供指导。

Comments Accepted at the AAAI 2026 Summer Symposium Series. This paper was presented at the ICLR Re-Align workshop under a different name, "Accelerating Adversarial Suffix Optimization via Continuous Relaxation and Activation-Guided Objectives"

详情

展开后加载摘要…

URL PDF HTML 收藏