发表机构
Pedeciba Informática; InCo, Facultad de Ingeniería, Universidad de la República; Departamento de Informática e Inteligencia Artificial, Universidad Católica del Uruguay(Pedeciba信息学研究所; 共和国大学工程学院计算研究所; 乌拉圭天主教大学信息与人工智能系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过XAI技术分析Prompt Guard 2的决策机制,发现其依赖多标记累积贡献,且显著性引导的扰动可翻转预测并实现越狱,揭示解释方法可能降低对抗性攻击成本。
AI 中文摘要
大型语言模型(LLMs)在生产系统中的部署日益广泛,这引发了对其遭受对抗性操纵(如提示注入和越狱攻击)的担忧。基于分类器的护栏(如Prompt Guard 2)被广泛用作抵御此类攻击的第一道防线,但其内部决策逻辑对防御者和攻击者而言在很大程度上是不透明的。本文提出了一项探索性案例研究,应用可解释人工智能(XAI)技术来分析Prompt Guard 2如何区分恶意提示与良性提示。我们进行了四项实验来实证探究这一问题。在Vanilla Gradient和SHAP归因的指导下,我们发现Prompt Guard 2的决策依赖于许多标记的累积贡献,而非少数主导标记,然而基于显著性引导的同义词替换和句子级释义可以翻转其预测,同时仅修改文本中适度比例的部分,在某些情况下能成功实现对底层LLM的越狱。数据集规模的显著性分析进一步表明,未被检测到的注入提示系统性地缺乏分类器所依赖的词汇标记。我们讨论了这些发现对基于分类器的护栏的设计与评估的启示,并论证旨在支持透明度的解释方法可能同时降低构建成功对抗性绕过的成本。
英文摘要
Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their internal decision logic is largely opaque to both defenders and attackers. This paper presents an exploratory case study that applies explainable artificial intelligence (XAI) techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts. We conduct four experiments to probe this question empirically. Guided by Vanilla Gradient and SHAP attributions, we find that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substitution and sentence-level paraphrasing can flip its predictions while altering only a moderate fraction of the text, in some cases yielding a successful jailbreak against the underlying LLM. A dataset-scale saliency analysis further shows that undetected injection prompts systematically lack the lexical markers the classifier relies on. We discuss the implications of these findings for the design and evaluation of classifier-based guardrails, and argue that explanation methods intended to support transparency can simultaneously lower the cost of constructing successful adversarial bypasses.