pikit:用于间接提示注入研究与评估的可组合工具包
pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation
浏览论文内容
中文总结 AI 辅助
本文提出pikit工具包,系统评估间接提示注入的攻击、渠道与防御,在真实环境中使攻击成功率降低71.8%,并实现可复现的自动记录。
中文摘要 AI 辅助
间接提示注入将恶意指令嵌入到由基于LLM的智能体检索的外部内容中,在未经用户授权的情况下改变目标行为。我们介绍了pikit,一个旨在系统评估这些威胁的研究工具包,涵盖三个核心维度:攻击(13种方法)、渠道(文本和文件模式下的16种载体)和防御(9种预防策略和3种离线检测基线)。基于装饰器注册表构建,pikit支持在不修改核心代码的情况下无缝扩展自定义组件,同时统一的craft() API可在单次调用中组合任意攻击和渠道。我们在由匿名LLM驱动的pi编码智能体上,在类似生产的环境中评估了该工具包。针对高风险攻击对9种预防策略进行基准测试,攻击成功率相对降低71.8%,其中few_shot_warning和instruction_hierarchy提供了最强的保护。离线检测基线实现了完美的精确率但召回率较低,表明启发式检测器补充而非替代提示级防御。为确保可复现性,每次运行自动记录完整提示、智能体事件轨迹、会话记录和判定记录。我们的代码可在https://this URL获取。
英文摘要
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 carriers across text and file modes), and defenses (9 prevention strategies and 3 offline detection baselines). Built on a decorator-based registry, pikit enables seamless extension of custom components without modifying core code, while a unified craft() API composes arbitrary attacks and channels in a single call. We evaluated the toolkit on the pi coding agent powered by an anonymized LLM in a production-like environment. Benchmarking 9 prevention strategies against high-risk attacks yields a 71.8\% relative reduction in attack success rate, with few\_shot\_warning and instruction\_hierarchy providing the strongest protection. Offline detection baselines achieve perfect precision but low recall, demonstrating that heuristic detectors complement rather than replace prompt-level defenses. To ensure reproducibility, each run automatically logs full prompts, agent event traces, session transcripts, and verdict records. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/pikit.
发表机构
- Zhuque Lab(朱雀实验室)
机构由 AI 辅助整理,请以论文原文为准。