发表机构
Johns Hopkins University; Washington University in St. Louis; University of Virginia; NVIDIA; Virginia Tech(约翰斯·霍普金斯大学; 圣路易斯华盛顿大学; 弗吉尼亚大学; 英伟达; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出自我反思微调(SRFT)框架,通过让智能体从对抗性失败经验中学习,增强其对提示注入攻击的鲁棒性,实验显示显著降低攻击成功率并保持任务性能。
AI 中文摘要
大型语言模型(LLM)智能体越来越多地部署在工具增强的环境中,但它们对外部输入的依赖使其极易受到提示注入攻击,这些攻击可能劫持任务目标。现有的安全对齐方法依赖于静态专家轨迹或偏好优化,限制了它们对自适应攻击模式的泛化能力。在这项工作中,我们提出了自我反思微调(SRFT),这是一种训练框架,使智能体能够通过从自身在对抗条件下的失败经验中学习来提高鲁棒性。SRFT不是被动地模仿专家行为,而是通过注入攻击构建的受损轨迹暴露给智能体,并利用专家模型生成结构化的自我反思推理,对比不安全行为与最优行为。这种反思性监督教会智能体识别恶意指令,推理其后果,并保持与原始用户意图的一致性。我们将该框架实例化为SR-Agent,基于Llama-3.1-8B-Instruct和Qwen3-8B构建,并在静态和自适应提示注入基准上进行了评估。实验结果表明,SRFT显著降低了攻击成功率,同时保持了任务性能,并在自适应攻击下表现出强大的泛化能力。这些发现表明,通过自我反思从失败中学习是构建鲁棒且安全的LLM智能体的一个有前景的方向。我们的代码已在此https URL发布。
英文摘要
Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely on static expert trajectories or preference optimization, limiting their ability to generalize to adaptive attack patterns. In this work, we propose Self-Reflection Fine-Tuning (SRFT), a training framework that enables agents to improve robustness by learning from their own failure experiences under adversarial conditions. Instead of passively imitating expert behaviors, SRFT exposes the agent to compromised trajectories constructed via injected attacks, and leverages an expert model to generate structured self-reflection reasoning that contrasts unsafe and optimal actions. This reflective supervision teaches the agent to identify malicious instructions, reason about their consequences, and maintain alignment with the original user intent. We instantiate this framework in SR-Agent, built on Llama-3.1-8B-Instruct and Qwen3-8B, and evaluate it on both static and adaptive prompt injection benchmarks. Experimental results show that SRFT substantially reduces attack success rates while preserving task performance, and demonstrates strong generalization under adaptive attacks. These findings suggest that learning from failure via self-reflection is a promising direction for building robust and secure LLM agents. Our code is released at https://github.com/Eden-Wang1710/srft-repo.