AIP:通过对抗性指令提示颠覆检索增强生成
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
- Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出AIP攻击,通过对抗性指令提示而非用户查询来操纵RAG检索行为,并利用多样化查询生成与遗传算法联合优化,在保持良性功能的同时实现高达95.23%的攻击成功率,揭示了共享指令提示中的关键安全漏洞。
AI中文摘要:
检索增强生成(RAG)通过从外部来源检索相关文档来增强大语言模型(LLMs),以提高事实准确性和可验证性。然而,这种依赖在检索流水线中引入了超出LLM本身的新攻击面。尽管先前的RAG攻击已经暴露了此类漏洞,但它们主要依赖于操纵用户查询,而在实践中由于用户输入固定或受保护,这通常不可行。这种狭窄的关注忽视了一个更现实且更隐蔽的向量:指令提示,它们被广泛复用、公开共享且很少被审计。其隐式信任使其成为对手隐蔽操纵RAG行为的有吸引力的目标。我们引入了一种针对对抗性指令提示(AIP)的新型攻击,该攻击利用对抗性指令提示通过微妙地改变检索行为来操纵RAG输出。通过将攻击面转移到指令提示上,AIP揭示了可信但看似良性的界面组件如何被武器化以降低系统完整性。该攻击被设计为实现三个目标:(1)自然性,以逃避用户检测;(2)实用性,以鼓励使用提示;(3)鲁棒性,以在多样化的查询变体中保持有效。我们提出了一种多样化查询生成策略,模拟用户查询中现实的语言变化,从而能够发现可跨释义和改写泛化的提示。在此基础上,开发了一种基于遗传算法的联合优化方法,通过平衡攻击成功率、干净任务实用性和隐蔽性来进化对抗性提示。实验结果表明,AIP在保持良性功能的同时实现了高达95.23%的ASR。这些发现揭示了RAG系统中一个关键且此前被忽视的漏洞,强调了重新评估共享指令提示的必要性。
英文摘要:
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by retrieving relevant documents from external sources to improve factual accuracy and verifiability. However, this reliance introduces new attack surfaces within the retrieval pipeline, beyond the LLM itself. While prior RAG attacks have exposed such vulnerabilities, they largely rely on manipulating user queries, which is often infeasible in practice due to fixed or protected user inputs. This narrow focus overlooks a more realistic and stealthy vector: instructional prompts, which are widely reused, publicly shared, and rarely audited. Their implicit trust makes them a compelling target for adversaries to manipulate RAG behavior covertly. We introduce a novel attack for Adversarial Instructional Prompt (AIP) that exploits adversarial instructional prompts to manipulate RAG outputs by subtly altering retrieval behavior. By shifting the attack surface to the instructional prompts, AIP reveals how trusted yet seemingly benign interface components can be weaponized to degrade system integrity. The attack is crafted to achieve three goals: (1) naturalness, to evade user detection; (2) utility, to encourage use of prompts; and (3) robustness, to remain effective across diverse query variations. We propose a diverse query generation strategy that simulates realistic linguistic variation in user queries, enabling the discovery of prompts that generalize across paraphrases and rephrasings. Building on this, a genetic algorithm-based joint optimization is developed to evolve adversarial prompts by balancing attack success, clean-task utility, and stealthiness. Experimental results show that AIP achieves up to 95.23% ASR while preserving benign functionality. These findings uncover a critical and previously overlooked vulnerability in RAG systems, emphasizing the need to reassess the shared instructional prompts.