发表机构
George Mason University; Virginia Tech; Missouri University of Science and Technology; Visa(乔治梅森大学; 弗吉尼亚理工大学; 密苏里科技大学; 维萨公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示护栏模型在重复输入下会发生MAL→BEN标签翻转(Overflip),影响多数轻量级模型,威胁大于传统注意力稀释,需长度鲁棒评估与缓解。
AI 中文摘要
护栏模型是部署在基于LLM的服务中,用于筛查恶意提示和响应的分类器。为了满足延迟约束,许多轻量级护栏采用紧凑的Transformer骨干网络(如DeBERTa),这些网络使用短上下文窗口(通常为512个token)进行训练,并依赖分桶相对位置编码来处理更长的输入。先前的评估假设护栏的决策在输入变长时是稳定的。我们证明这一假设可能失效。我们识别出Overflip,一种重复诱导的不稳定性,其中重复提示会导致护栏的预测在序列增长时发生翻转(MAL→BEN)。我们在9个广泛使用的轻量级护栏模型上进行了实验。其中5个在100个提示的基准上表现出MAL→BEN翻转,置信度边际随重复而稳步缩小。在这些易受影响的模型中,翻转率从8%到92%不等,首次翻转发生在约2.6k至9.4k个token处。我们的分析表明,Overflip不同于传统的注意力稀释基线,后者旨在将模型的注意力从与恶意内容相关的token上转移开,而Overflip则将其转向无关内容,如良性填充或打乱。虽然Overflip保留了恶意内容,但它使token级注意力在重复结构上均匀化,并诱导出与填充不同的、更渐进的注意力分散轨迹。此外,Overflip对LLM服务构成的威胁比传统注意力稀释方法更大。因为被绕过的提示在语义上保持完整,并且仍能被下游业务LLM轻易理解,它可以在通过护栏后传递恶意意图。这些发现暴露了重复作为护栏模型的攻击面,并促使长度鲁棒的评估和缓解措施。
英文摘要
Guardrail models are classifiers deployed to screen malicious prompts and responses in LLM-based services. To meet latency constraints, many lightweight guardrails adopt compact Transformer backbones (e.g., DeBERTa) that are trained with short context windows (typically 512 tokens) and rely on bucketed relative positional encodings to process longer inputs. Prior evaluations assume that a guardrail's decision is stable as the input is lengthened. We show that this assumption can fail. We identify Overflip, a repetition-induced instability where repeating a prompt causes the guardrail's prediction to flip (MAL$\to$BEN) as the sequence grows. We conduct experiments on 9 widely used lightweight guardrail models. Five exhibit MAL$\to$BEN flips on a benchmark of 100 prompts, with confidence margins shrinking steadily with repetition. Among these vulnerable models, flip rates range from 8% to 92%, with first flips occurring at roughly 2.6k--9.4k tokens. Our analysis suggests Overflip differs from traditional attention-dilution baselines, which aim to divert the model's attention away from tokens associated with malicious content, shifting it instead toward unrelated content, such as benign padding or shuffling. While Overflip preserves malicious content, it homogenizes token-level attention over repeated structure and induces a distinct, more gradual attention-dispersion trajectory than padding. Moreover, Overflip poses a greater threat to LLM services than traditional attention dilution methods. Because the bypassed prompt remains semantically intact and is still readily understood by downstream business LLMs, it can transmit malicious intent after passing the guardrail. These findings expose repetition as an attack surface for guardrail models and motivate length-robust evaluation and mitigation.
Comments12 pages, 5 figures