arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03693cs.CRcs.AI

AlcaTRAz——针对越狱攻击的锚定树规则防御

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

  • Brno University of Technology(布尔诺理工大学)
  • Red Hat(红帽公司)

机构由 AI 辅助整理,请以论文原文为准。

Jakub Reš, Petr Kaška, Martin Perešíni, Martin Ukrop, Kamil Malinka

AI总结:

该研究提出仅需输入文本的提示级防御方法AlcaTRAz,在33个模型、22种攻击的测试中,多数场景下综合性能优于三种基线,可显著降低越狱成功率且对良性查询影响极小。

AI中文摘要:

大型语言模型(LLMs)易受越狱攻击影响,此类攻击通过精心设计的提示词绕过安全对齐。许多现有防御措施需要访问模型权重或内部组件,难以应用于黑盒部署场景。本文提出AlcaTRAz(Anchored Tree-Rule defense Against jailbreaks,针对越狱攻击的锚定树规则防御),这是一种仅基于输入文本运行的提示级防御,无需修改或重新训练目标模型。该方法自动学习可迁移的转换规则,在选定位置插入可控的字符级扰动,从而破坏越狱攻击利用的结构规律性,同时在很大程度上保留模型对良性查询的效用。我们在33个开放权重模型、22种越狱攻击类型,以及一个包含简短单轮良性问题的基准测试中对所提方法进行评估,与三个代表性提示级基线(Llama Guard、RA-LLM、Goal Prioritization)进行对比。在所有对比防御措施中,AlcaTRAz在73.4%的模型-攻击组合中取得了最佳的综合安全性与功能评分;在未防御场景中,聚合评分的众数为10(对恶意请求的最高严重程度响应),经防御后该众数变为2(近乎弃权(不执行)),同时良性查询的平均评分与未防御基线的差值在0.27分以内(0-10分制下,未防御基线为8.62,防御后为8.35)。AlcaTRAz大幅降低了但未完全消除越狱成功率:仍存在高严重程度的尾部情况,且本文未考虑自适应攻击者,因此将其定位为纵深防御策略中的一层,而非独立的安全保障。

英文摘要:

Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model. The method automatically learns a transferable transformation rule that inserts controlled character-level perturbations at selected positions, thereby disrupting structural regularities exploited by jailbreak attacks while largely preserving the model's utility on benign queries. We evaluate the proposed method across 33 open-weight models, 22 jailbreak attack types, and a benchmark of short, single-turn benign questions, comparing against three representative prompt-level baselines (Llama Guard, RA-LLM, Goal Prioritization). Among the compared defenses, AlcaTRAz achieves the best composite security and functionality score in 73.4 % of model-attack combinations and shifts the aggregate score from a modal value of 10 (maximal-severity response to the malicious request) in the undefended setting to a modal value of 2 (near-refusal) after defense, while keeping the mean benign score within 0.27 points of the undefended baseline (8.35 vs. 8.62 on a 0-10 scale). AlcaTRAz substantially reduces but does not eliminate jailbreak success: a high-severity tail remains, and we do not consider adaptive attackers, so we position it as one layer within a defense-in-depth strategy rather than a standalone guarantee.

补充信息

↑