arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32602cs.LGcs.AI

AnchorRep:通过表示排斥防御大语言模型的跨模型对抗性迁移

AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion

Gal Wertheizer, Rom Himelstein, Tomer Peretz, Avi Mendelson

首次发表
浏览论文内容

中文总结 AI 辅助

AnchorRep通过轻量级LoRA适配器排斥有害提示的表示,防御跨模型对抗性迁移,将攻击成功率降至≤1.1%,并引入良性乱码率评估退化输出。

中文摘要 AI 辅助

在单个开源权重大语言模型上优化的对抗性攻击可以迁移到并越狱架构不同的模型,使得攻击者能够利用对一个模型的白盒访问权限来破坏独立部署的系统。这造成了模型间的共享脆弱性,然而现有防御措施并非针对这种跨模型威胁而设计。我们发现跨模型迁移与共享的内部表示几何结构相一致,使其成为自然的防御目标。AnchorRep直接针对这一几何结构,使用一个轻量级LoRA适配器,将受防御模型对有害提示的内部表示推离同一提示下冻结锚模型的表示。训练使用少量有害提示且无需对抗性示例。在五个模型和四个架构族中,AnchorRep将跨模型攻击成功率在2000次迁移攻击上降至≤1.1%(其中两个模型为0%),包括在Mistral上的最大降幅(36%→1.1%)。现有防御可以降低迁移,但代价高昂,要么诱导高达77%的退化良性输出,要么将过度拒绝率提高多达18%。由于这种退化的良性输出未被基于拒绝的标准指标捕获,我们引入了良性乱码率来量化它们。我们的结果表明,通过塑造表示几何结构可以实现跨模型鲁棒性,而无需依赖特定于攻击的训练。

英文摘要

Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to <=1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% -> 1.1%). Existing defenses can reduce transfer, but only at high cost either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training

补充信息

↑