RAG-PIBench:面向可信RAG系统中提示注入检测的泄漏感知基准
RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
浏览论文内容
中文总结 AI 辅助
针对RAG系统易受提示注入攻击的问题,提出泄漏感知基准RAG-PIBench(含4,876个示例),比较多种检测器,DistilBERT最优(F1=0.896),验证了泄漏感知设计与稀疏基线的有效性。
中文摘要 AI 辅助
检索增强生成(RAG)系统容易受到嵌入在检索内容中的提示注入攻击。我们引入了RAG-PIBench,一个用于RAG风格提示注入检测的基准,包含4,876个上下文示例,分为固定的训练集、验证集和受保护测试集。通过采用泄漏感知的构建流程和严格的评估协议,我们比较了基于关键词、语义参考、TF-IDF和基于Transformer的检测器。DistilBERT在受保护测试集上取得了最佳性能(F1 = 0.896,PR-AUC = 0.968),而TF-IDF SVM和逻辑回归仍具有竞争力。我们的结果证明了泄漏感知基准设计和强稀疏基线在RAG系统中实现可靠提示注入检测的价值。
英文摘要
Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
发表机构
- Birzeit University(比尔泽特大学)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。