发表机构
Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Persist-3,一种基于逐源共形p值的序贯检测方法,用于从攻击提示流中识别未知越狱攻击源,理论保证检测概率随窗口指数趋近于1,并在MTK和Prompt Guard 2上显著增强防御。
AI 中文摘要
针对大型语言模型的越狱防御通常是在一组固定的已知攻击方法上校准的。然而,部署后新的攻击方法不断出现,现有检测器在新攻击下性能会下降(Piet等人,2025)。本文研究如何从攻击提示流中检测出提示来自已知攻击源之外的攻击源。我们将其形式化为一个序贯检验,其零假设是提示流来自某个已知攻击源。在方法论上,我们针对每个已知攻击源计算一个共形p值,并且仅当针对每个已知攻击源的证据都足够强时才拒绝零假设。利用这些逐源p值,我们构建了Persist-3,并分析了其平稳误报概率和检测能力。Persist-3的技术贡献在于其检测保证仅使用最近观测的有限窗口。我们的理论分析表明,随着窗口增大,其检测概率以指数速度趋近于1,因此一个短窗口就足以检测出与所有已知攻击源分离良好的未见攻击源。在实验上,我们在两种最新的越狱防御方法MTK(Zhang等人,2026)和Prompt Guard 2(Chennabasappa等人,2025)之上验证了Persist-3,并观察到显著的防御增强效果。
英文摘要
Jailbreak defenses for large language models are usually calibrated on a fixed set of known attack methods. However, new attack methods keep appearing after deployment, and existing detectors are observed to degrade under the new attacks (Piet et al., 2025). This paper studies how to detect, from a stream of attack prompts, that the prompts come from an attack source outside the known ones. We formulate it as a sequential test whose null hypothesis is that the stream comes from one of the known sources. Methodology-wise, we compute a conformal p-value against each known source, and reject the null only when the evidence against every known source is large. Using these source-wise p-values, we build Persist-3 and analyze its stationary false-alarm probability and detection power. The technical contribution of Persist-3 lies in a detection guarantee that uses only a finite window of recent observations. Our theoretical analysis reveals that its detection probability approaches one exponentially fast as the window grows, and thus a short window suffices to detect unseen sources that are well separated from all known sources. Empirically, we validate Persist-3 on top of two recent jailbreak defenses, MTK (Zhang et al., 2026) and Prompt Guard 2 (Chennabasappa et al., 2025), and observe a significant defense enhancement.