arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10742cs.CRcs.AI

BRANCH:绕过多扫描器AI护栏

BRANCH: Bypassing Multi-Scanner AI Guardrails

William Hackett, Peter Garraghan

首次发表
浏览论文内容

中文总结 AI 辅助

针对多扫描器AI护栏系统,提出BRANCH分支树搜索绕过方法,可在减少查询量与时间的同时实现100%攻击成功率,且绕过可迁移至多数未见过的护栏。

中文摘要 AI 辅助

AI系统越来越依赖大型语言模型(LLMs)作为核心推理引擎,使其成为提示注入和越狱攻击的目标。护栏(guardrails)会监控并验证模型的输入和输出,但它们孤立的、面向特定任务的检测方式在分类上存在漏洞,容易被绕过。为此,由多个扫描器组成的护栏系统应运而生,这些扫描器协同检测不同类型的恶意指令,而分类边界间共享的潜在表征使得已有的绕过技术失效。我们提出BRANCH,一种针对多扫描器护栏系统的绕过方法。该方法利用分支树搜索方法,动态对单个扫描器施加对抗扰动,随后基于所有护栏系统扫描器的整体改进进行扰动优化和技术选择,有效将绕过评估与攻击信号优化解耦。研究结果表明,BRANCH在120种场景下对6个护栏系统实现了100%的攻击成功率,与现有技术相比,查询量减少72%, wallclock时间缩短4.5倍,同时保留了绕过内容的语义。我们还发现,BRANCH生成的绕过可迁移至29个未见过的护栏,包括8个商业黑盒护栏,在无需额外优化的情况下,部分场景的攻击成功率提升至100%。

英文摘要

AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective. We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems. Our method leverages a branching tree search approach that dynamically applies adversarial perturbation against individual scanners, with subsequent perturbation optimization and technique selection based on overall improvement across all guardrail system scanners, effectively decoupling bypass evaluation from attack signal optimization. Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass. We also show how bypasses generated by BRANCH transfer to 29 unseen guardrails, including 8 commercial black-box guardrails, improving attack success in some cases up to 100% with no additional optimization.

发表机构

  • Mindgard
  • Lancaster University(兰卡斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑