发表机构
Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM易受间接提示注入攻击的问题,提出上下文感知的恶意句子分类方法及两种对抗训练技术,在静态和自适应攻击下均提升了IPI防御性能,且需根据领域调整参数。
AI 中文摘要
现代大型语言模型(LLM)出色的指令遵循能力使其可作为智能体的核心,用于自主完成日益复杂的任务,但这也使其易受到嵌入恶意指令的文本攻击,常见变体称为间接提示注入(IPI)。解决此漏洞的基础任务是将给定文本成功分割为良性和恶意句子(若存在)。尽管已提出多种针对该任务的方法,但尚无检测器能在片段级别结合查询相关检测,且无法抵御智能体执行中可实现的自适应规避攻击。我们通过开发一种同时具备上下文感知和查询感知的恶意句子分类方法,解决了前者的局限;为使分类器具备抵御规避攻击的能力,我们提出两种对抗训练方法:第一种是直接适配的特征空间对抗训练(AT),其中通过嵌入空间中基于投影梯度的优化来近似规避攻击;第二种是在AT循环中通过基于LLM的paraphrasing模拟可实现的规避攻击。关键在于,我们对两种AT变体进行参数化,以在实用性和攻击鲁棒性之间实现平稳权衡。在使用间接提示注入基准进行的大量实验中,我们表明,所提方法在静态攻击下优于最先进的IPI防御基线;而在自适应攻击下,我们的AT变体提供了显著更高的实用性、更低的攻击成功率,且常同时实现这两点。最后,我们发现最佳AT参数可能与特定应用领域密切相关,因此,在实践中可能需要对恶意文本检测器进行依赖领域的调整。我们的代码公开于此https://URL。
英文摘要
The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks. Therein, however, also lies their vulnerability to attacks which embed malicious instructions in text, common variants of which are known as indirect prompt injection (IPI). A fundamental task in addressing this vulnerability is successful segmentation of a given text into benign and malicious sentences (if any). While a number of approaches for this task have been proposed, no detector combines query-relative detection at the segment level, and none are hardened against adaptive evasion attacks realizable in agentic executions. We address the former limitation by developing an approach for malicious sentence classification that is both context- and query-aware. Next, to harden the resulting classifier against evasion, we present two adversarial training methods. The first is directly adapted feature-space adversarial training (AT) in which evasions are approximated using projected-gradient-based optimization in the embedding space. The second simulates realizable evasion attacks in the AT loop through LLM-based paraphrasing. Crucially, we parametrize both AT variants to facilitate a smooth tradeoff between utility and attack robustness. In extensive experiments using indirect prompt injection benchmarks we show that the proposed approach outperforms state-of-the-art IPI defense baselines under static attacks, while in the case of adaptive attacks, our AT variants provide significantly higher utility, lower attack success rate, and often both. Finally, we show that the best AT parameters can depend intimately on the particular application domain. Consequently, domain-dependent tuning of malicious text detectors is likely necessary in practice. Our code is publicly available at https://github.com/tavia-liu/CAD.