AI 中文总结
针对LLMs防御间接提示注入时的安全-效用权衡问题,提出AEGIS模型,通过指令敏感投影器与统一多层共识机制实现出色的攻击检测性能,缓解IPI攻击。
AI 中文摘要
大型语言模型(LLMs)已被集成到复杂生态系统(如代码智能体)中,而间接提示注入(IPI)攻击已成为其安全部署的关键障碍。攻击者利用LLMs无法区分“指令”与“数据”的特性,通过恶意注入的指令操纵LLMs。然而,现有防御面临难以解决的安全-效用权衡问题:大多数防护措施要么延迟较高,要么存在严重的过度弃权(不执行)问题。在本文中,我们首先通过理论和实证证据证明LLMs本质上能够将指令与数据分离。受此见解启发,我们提出AEGIS(Adaptive Ensemble Guard for Injection Shielding,即用于注入防护的自适应集成防护装置)。AEGIS提取指令敏感投影器以识别恶意指令,并利用统一多层共识机制聚合网络深度中拓扑不同的信号。实证评估表明,与基线方法相比,AEGIS对启发式攻击和基于优化的攻击均实现了出色的检测性能,凸显了其缓解IPI攻击的潜力。代码可在该https URL获取。
英文摘要
Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to their safe deployment. Attackers exploit LLMs' indistinguishability between "instructions" and "data" to manipulate LLMs via maliciously injected instructions. Existing defenses, however, face an intractable safety-utility trade-off: most guardrails either incur high latency or suffer from severe over-refusal. In this paper, we first demonstrate that LLMs can separate instruction from data intrinsically with both theoretical and empirical evidence. Inspired by this insight, we propose AEGIS (Adaptive Ensemble Guard for Injection Shielding). AEGIS extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth. Empirical evaluations show that AEGIS achieves remarkable detection performance against both heuristic and optimization-based attacks compared to baselines, highlighting its potential to mitigate IPI. Code is available at https://github.com/xaddwell/AEGIS
CommentsExtended Content of EMNLP 2026 Findings