AI 中文总结
本研究提出PIMiner智能体系统,用于自动提示注入红队测试,该系统构建可迁移策略库,仅需少量查询即可在IPIArena、AgentDojo基准上对多款LLM实现高攻击成功率,解决现有方法泛化性差的问题。
AI 中文摘要
提示注入对大语言模型(LLM)智能体构成重大安全风险,因此高效且有效的红队测试对于评估此类风险及收集训练数据以改进防御措施至关重要。现有最先进的提示注入红队测试方法主要依赖强化学习(RL),生成的攻击者模型往往难以泛化到新的目标LLM。本研究开发了PIMiner,一种用于提示注入红队测试的智能体系统。训练期间,PIMiner在一系列(数据集,目标模型)对上进行训练,并从头构建策略库;测试时,学习到的策略库可直接迁移至先前未见过的目标LLM,无需额外训练,且每个测试样本仅需向目标智能体发送少量查询(例如10次)。实验结果表明,PIMiner表现出强劲性能:在IPIArena上,其对Gemini-2.5-Pro的攻击成功率(ASR)达76.2%,对GPT-5.1达61.9%,对Claude-Sonnet-4.5达42.9%;在AgentDojo上,其对Gemini-2.5-Pro的ASR达86.7%,对GPT-5.1达53.3%,对Claude-Sonnet-4.5达40.0%。
英文摘要
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PIMiner is trained on a sequence of (dataset, target model) pairs and builds a strategy library from scratch. At test time, the learned strategy library can be directly transferred to a previously unseen target LLM without additional training. PIMiner requires only a small number of queries to a target agent (e.g., 10) per test sample. Experimental results demonstrate that PIMiner achieves strong performance. On IPIArena, it attains a 76.2% ASR against Gemini-2.5-Pro, 61.9% ASR against GPT-5.1, and 42.9% ASR against Claude-Sonnet-4.5. On AgentDojo, it achieves an 86.7% ASR against Gemini-2.5-Pro, 53.3% ASR against GPT-5.1, and 40.0% ASR against Claude-Sonnet-4.5.
CommentsOur code is available at https://github.com/wang-yanting/PIMiner