发表机构
École polytechnique; Institut Polytechnique de Paris; CNRS; Google DeepMind; Agence Ministérielle pour l’IA de Défense; IRT SystemX(巴黎综合理工学院; 巴黎理工学院; 法国国家科学研究中心; 谷歌DeepMind; 法国国防部人工智能局; 法国系统X技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具型语言模型智能体的间接提示注入攻击,提出RAISED训练框架,结合自生成与自蒸馏,在显著降低攻击成功率的同时保持智能体及通用基准上的效用。
AI 中文摘要
使用工具的语言模型智能体容易受到间接提示注入的攻击,因为它们必须对不可信的外部内容采取行动。现有的训练时防御可以降低攻击成功率,但往往以牺牲通用能力为代价。我们表明,基于训练的防御会诱导模型输出分布发生显著漂移,即使在良性设置下也会改变其行为,并为效用退化提供了潜在机制。我们进一步识别了这些防御的一个失败模式:在良性工具使用任务中,模型会回避完成授权任务所需的步骤,特别是当该步骤由工具输出指示时。为了解决这些局限性,我们引入了RAISED(通过自蒸馏实现鲁棒攻击不变性),这是一种结合自生成和自蒸馏的训练框架。模型首先生成自己的工具使用场景,重点放在任务完成需要根据工具输出的合法指导采取行动的情况。然后,通过自蒸馏,学生模型被训练为在相同轨迹的干净和注入变体上匹配教师模型的干净上下文行为。RAISED显著降低了工具响应中提示注入的攻击成功率,同时与先前的基于训练的防御不同,它在智能体和通用基准上保持了效用。
英文摘要
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.