CrackedPDFs:用于PDF中隐藏提示注入的受控基准测试
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
浏览论文内容
中文总结 AI 辅助
研究针对PDF中隐藏提示注入问题,引入CrackedPDFs基准测试,评估多种检测方法,如PromptGuard、结构学习模型等,结果显示文档感知混合检测在受控配对评估下有效,但缺乏现实世界鲁棒性和跨家族泛化能力。
中文摘要 AI 辅助
基于文档的大型语言模型(LLM)系统在护栏检查之前通常会将PDF扁平化,这一步骤可能会丢弃用户从未看到过指令的证据。我们引入了CrackedPDFs,这是一个用于PDF中隐藏提示注入的受控基准测试。该基准测试包含来自4983个基础文档生成的29322个PDF,包括9774个注入文件和19548个良性或匹配混淆文件。我们评估了PromptGuard和规则基线,还评估了仅基于结构学习的模型和经过净化的混合检测器。评估使用了留出的来源分割和配对的良性混淆控制,还使用了标签洗牌检查和捷径审计。在一个2919个文档的留出测试集上,混合检测器的F1值达到0.960,ROC-AUC为0.998,PR-AUC为0.997。它在973对中的95.9%中,将注入文件排在匹配的良性混淆文件之上。仅给出提取的文本时,PromptGuard召回率较低。仅基于结构学习的模型在配对控制下较弱。仅文本的TF-IDF模型在留出测试中获得完美分数,但未能通过捷径审计。这些结果表明,在受控配对评估下,文档感知混合检测是有用的,但它们并未显示出广泛的现实世界鲁棒性或可靠的跨家族泛化能力。
英文摘要
Document-based LLM systems often flatten a PDF before guardrails inspect it. That step can discard evidence that an instruction was never visible to the user. We introduce CrackedPDFs, a controlled benchmark for hidden prompt injection in PDFs. The benchmark contains 29,322 generated PDFs from 4,983 base documents. It includes 9,774 injected files and 19,548 benign or matched-confounder files. We evaluate PromptGuard and a rule baseline. We also evaluate structural-only learned models and a sanitized hybrid detector. The evaluation uses held-out provenance splits and paired benign-confounder controls. It also uses label-shuffle checks and shortcut audits. On a 2,919-document held-out test set, the hybrid detector reaches 0.960 F1. ROC-AUC is 0.998 and PR-AUC is 0.997. On a balanced subset containing 973 injected PDFs and 973 matched benign confounders, the hybrid detector achieves 95.9% classification accuracy. Using score ordering, it ranks the injected member above its matched confounder in 100% of 973 pairs. PromptGuard has low recall when given extracted text only. Structural-only learned models are weak under paired controls. A text-only TF-IDF model reaches perfect held-out scores but fails shortcut audits. These results show that document-aware hybrid detection is useful under controlled paired evaluation. They do not show broad real-world robustness or reliable cross-family generalization.
发表机构
- University of California, Berkeley(加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。