用于大语言模型越狱的可解释 N-gram 困惑度威胁模型
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
- University of Tübingen(图宾根大学)
- Tübingen AI Center(图宾根人工智能中心)
- Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
- ELLIS Institute Tübingen(图宾根ELLIS研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出基于 1T tokens N-gram 语言模型的可解释、LLM 无关威胁模型,用以统一评估越狱攻击,发现离散优化攻击更强且有效攻击依赖罕见 bigram。
AI中文摘要:
大量越狱攻击已被提出,用于从经过安全调优的 LLM 获取有害回复。这些方法在其原始设置中大多能成功胁迫目标输出,但其攻击在流畅度和计算开销方面差异显著。本文提出一个统一威胁模型,以对这些方法进行有原则的比较。该威胁模型检查给定越狱是否可能出现在文本分布中。为此,我们在 1T tokens 上构建了一个 N-gram 语言模型;与基于模型的困惑度不同,它支持与 LLM 无关、非参数且本质可解释的评估。我们将流行攻击适配到该威胁模型,并首次利用它在同等基础上对这些攻击进行基准测试。经过广泛比较,我们发现针对现代安全调优模型的攻击成功率低于此前报告,且基于离散优化的攻击显著优于近期基于 LLM 的攻击。由于本质上可解释,该威胁模型支持对越狱攻击进行全面分析和比较。我们发现,有效攻击会利用并滥用不常见的 bigram:要么选择真实世界文本中不存在的 bigram,要么选择罕见 bigram,例如特定于 Reddit 或代码数据集的 bigram。
英文摘要:
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but their attacks vary substantially in fluency and computational effort. In this work, we propose a unified threat model for the principled comparison of these methods. Our threat model checks if a given jailbreak is likely to occur in the distribution of text. For this, we build an N-gram language model on 1T tokens, which, unlike model-based perplexity, allows for an LLM-agnostic, nonparametric, and inherently interpretable evaluation. We adapt popular attacks to this threat model, and, for the first time, benchmark these attacks on equal footing with it. After an extensive comparison, we find attack success rates against safety-tuned modern models to be lower than previously presented and that attacks based on discrete optimization significantly outperform recent LLM-based attacks. Being inherently interpretable, our threat model allows for a comprehensive analysis and comparison of jailbreak attacks. We find that effective attacks exploit and abuse infrequent bigrams, either selecting the ones absent from real-world text or rare ones, e.g., specific to Reddit or code datasets.