arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05850cs.LGcs.AIcs.CR

SAFEGuard:通过有害语义分析与流畅度测量检测基于优化的越狱攻击

SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

Quoc Viet Vo, Trung Le, Damith C. Ranasinghe, Ehsan Abbasnejad

首次发表
浏览论文内容

中文总结 AI 辅助

SAFEGuard通过混合流畅度测量和有害语义分析,统一检测基于优化的越狱攻击,显著提升检测准确性。

中文摘要 AI 辅助

尽管在使大型语言模型(LLMs)与人类价值观对齐并确保安全部署方面付出了巨大努力,但最近的研究表明,LLMs仍然容易受到对抗性越狱攻击,这些攻击可以绕过安全护栏并引发有害响应。许多防御方法被提出来检测越狱,但它们在应对广泛的基于优化的越狱机制方面效果有限,这些机制可以产生高度流畅优化或有害语义混淆的提示。为了解决这一挑战,我们提出了一个统一的检测框架SAFEGuard,它结合了基于跨层分布距离和困惑度的混合流畅度测量,以及通过梯度匹配进行的有害语义分析。我们的方法基于一个重要的观察:高流畅度提示将其恶意意图保持在接近有害提示的水平,而有毒语义混淆的提示通常注入乱码令牌序列。我们的评估表明,SAFEGuard在不同基于优化的越狱攻击中始终优于最先进的基线,并在准确性方面取得了显著提升。这强调了SAFEGuard在应对不断演变的越狱攻击方面的有效性。

英文摘要

Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to detect jailbreaks but they are limited in their effectiveness to counter wide-range optimization-based jailbreak mechanisms that can yield highly fluency-optimized or harmful semantic obfuscated prompts. To tackle this challenge, we propose a unified detection framework SAFEGuard which incorporates a hybrid fluency measurement based on cross-layer distribution distance and perplexity, and the analysis of harmful semantics through gradient matching. Our method is grounded in a paramount observation: high fluency prompts maintain their malicious intention close to harmful prompts while harmful semantic obfuscated prompts often inject gibberish token sequences. Our evaluation demonstrates that SAFEGuard consistently outperforms state-of-the-art baselines and achieves significant improvement in accuracy across different optimization-based jailbreaks. This underscores the effectiveness of SAFEGuard against evolving jailbreak attacks.

发表机构

  • Australian Institute for Machine Learning(澳大利亚机器学习研究所)
  • Adelaide University(阿德莱德大学)
  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑