arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当大语言模型防御适得其反时:刻画安全性、性能和成本的权衡

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Tong Zhang, Zexin Li, Simin Chen, Yun Peng

arXiv 2607.24392首次发表:更新:

发表机构

Fudan University; University of California, Riverside; Columbia University(复旦大学; 加州大学河滨分校; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型越狱防御的安全性、性能和成本权衡,按操作策略组织防御并研究其副作用,发现不同防御在安全与可用性、效率权衡上有差异,为评估防御副作用及选择防御提供基准和指导。

AI 中文摘要

越狱防御对于保护大语言模型至关重要,但它们也可能带来削弱模型效用的次要成本。我们沿着性能影响、对良性输入的过度拒绝和推理成本这三个维度对这些防御权衡进行了系统研究。我们按操作策略对防御进行组织,而非将其视为单一类别,并研究不同策略如何与不同的副作用特征相关联。在跨最先进的防御方法、广泛使用的基准数据集和有代表性的开源大语言模型的研究中,我们发现防御很少能提高下游能力,而是在安全收益与可用性和效率的权衡方式上存在差异。特别是,基于规则的防御最能保持任务性能,高度保守的自我反思防御往往会增加过度拒绝,而多轮防御会产生最大的运行时开销。这些结果既为评估防御副作用提供了基准,也为在部署约束下选择防御提供了实际指导。

英文摘要

Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑