arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PIDS-Bench:在过度防御、混淆和分布偏移下评估提示注入检测器

PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift

Yusuf Khalid Shire, Sang-Chul Kim

arXiv 2609.15017首次发表:更新:

发表机构

Kookmin University(国民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PIDS-Bench通过多轴基准揭示提示注入检测器在分布偏移下存在来源敏感过度防御,即使高F1也无法保证低误报。

AI 中文摘要

提示注入检测器通常使用分布内测试数据上的聚合F1分数进行评估,这在对分布偏移下的行为提供的洞察有限,尤其是在决策边界的良性一侧,其中误报会带来直接的操作成本,却很少被测量。我们提出了PIDS-Bench,一个冻结的多轴基准,在固定阈值下联合评估攻击检测和良性误报行为,涵盖分布内输入、模仿注入结构但无恶意意图的硬良性提示、混淆攻击,以及领域和结构分布偏移。我们评估了七个检测器(学习基线、外部提示注入分类器和广泛安全比较器),以及一个基于规则的参考下限。多轴评估暴露了聚合F1所掩盖的失败模式。一个在留出分割上超过F1=0.98的检测器,仍然将来自公共语料库且仅限于安全相关内容的约三分之一的外部来源良性子集错误分类。在完整的阈值扫描和五个训练种子中,没有内部检测器达到同时满足F1>=0.95和硬良性FPR<=0.10的工作点。按来源分解,我们发现硬负样本增强几乎消除了策划压力输入上的过度防御,但在外部来源的提示上基本保持完整,我们将这种模式称为来源敏感过度防御。这种不对称性在两种微调架构中都存在,并且不会随着增强池的增长而减弱,外部来源的FPR仍然远高于0.10的目标。与外部来源分布匹配的增强是否能缩小这一差距尚未测试;仅阈值校准和策划式增强无法做到。

英文摘要

Prompt-injection detectors are typically evaluated using aggregate F1 on in-distribution test data, which offers limited insight into behavior under distribution shift, particularly on the benign side of the decision boundary, where false positives impose direct operational cost yet are seldom measured. We present PIDS-Bench, a frozen multi-axis benchmark that jointly evaluates attack detection and benign false-positive behavior at fixed thresholds, spanning in-distribution inputs, hard-benign prompts that mimic injection structure without malicious intent, obfuscated attacks, and domain and structural distribution shifts. We evaluate seven detectors (learned baselines, external prompt-injection classifiers, and broad-safety comparators) alongside a rule-based lower-bound reference. Multi-axis evaluation exposes a failure mode that aggregate F1 conceals. A detector exceeding F1 = 0.98 on the held-out split still misclassifies roughly one-third of an externally-sourced benign subset drawn from public corpora and restricted to security-adjacent content. Across a full threshold sweep and five training seeds, no internal detector reaches an operating point satisfying F1 >= 0.95 and hard-benign FPR <= 0.10 together on this stress distribution. Decomposing by provenance, we find that hard-negative augmentation nearly eliminates over-defense on curated stress inputs but leaves it substantially intact on externally-sourced prompts, a pattern we term provenance-sensitive over-defense. The asymmetry holds across both fine-tuned architectures and does not diminish as the augmentation pool grows, with the externally-sourced FPR remaining far above the 0.10 target. Whether augmentation matched to the externally-sourced distribution would close this gap is untested; threshold calibration and curated-style augmentation alone do not.

Comments21 pages, 3 figures, 18 tables

Journal refIEEE Access, vol. 14, pp. 134184-134205, 2026

DOI:10.1109/ACCESS.2026.3728186

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑