arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CASCADE 对抗越狱:跨阶段的组合与受控攻防评估

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Jiale Luo, Eric Han

arXiv 2609.21793首次发表:更新:

发表机构

School of Computing; National University of Singapore(计算学院; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM越狱防御在不同流水线阶段部署不清的问题,首次系统研究阶段内与跨阶段的防御组合,在统一威胁模型下标准化评估,发现无单一防御最优,但精选组合可在最小化效用降级下显著提升安全性。

AI 中文摘要

针对大型语言模型(LLM)的越狱攻击防御措施作用于不同的流水线阶段,例如输入修改或输出防护,但尚不清楚在每个阶段应部署哪些防御措施以及如何组合它们。先前的实证研究因攻击成功率定义和实验设置不一致而碎片化,大多在孤立状态下评估防御措施。在此,我们提出了据我们所知首次对流水线阶段内部和跨阶段的防御组合进行的系统性研究,采用一致的威胁模型,即直接、黑盒、单轮攻击。我们的决策框架通过带有受控查询预算的原则性攻击成功率公式以及明确的公平性规则,使评估标准化。在19种攻击和15种防御措施中,我们发现没有单一防御措施普遍最优,但精心选择的组合能在最小化效用降级的同时实现显著的安全性,为分层防御流水线提供了实用建议。

英文摘要

Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we present the first systematic study, to our knowledge, of defense combinations both within and across pipeline stages, under a consistent threat model of direct, black-box, single-turn attacks. Our decision framework standardizes evaluation through a principled attack-success-rate formulation with controlled query budgets, together with explicit fairness rules. Across 19 attacks and 15 defenses, we find that no single defense is universally best, but well-chosen combinations achieve substantial safety with minimal utility degradation, yielding practical recommendations for layered defense pipelines.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑