AI 中文总结
该研究提出COPA框架,将提示注入防御视为终身学习问题,通过GRPO优化和边际加权经验回放实现持续防御,使攻击成功率最高降低6.3倍、平均降低4.4倍。
AI 中文摘要
大型语言模型(LLMs)仍然容易受到提示注入攻击,这类攻击中嵌入在用户输入或外部内容中的对抗性指令会操纵模型行为并绕过安全措施。现有的防御措施大多是静态的,依赖于固定的对齐目标或针对特定攻击的过滤机制,当新的攻击策略出现时需要重新设计。虽然近期的终身对齐方法解决了用户偏好的变化问题,但它们没有考虑到会不断进化以利用先前学习的防御弱点的自适应对手。这种限制在实际部署中尤为重要,因为不断演变的攻击分布需要持续适应,同时不能牺牲对之前遇到的威胁的鲁棒性。我们提出了COPA,一个将提示注入防御视为终身学习问题的持续偏好优化框架。COPA不是一次性对齐,而是通过基于GRPO的优化逐步整合来自新观察到的攻击的反馈,并使用边际加权经验回放来保留对先前攻击类别的防御。这使得能够持续适应新出现的威胁,同时减轻灾难性遗忘并保留通用模型能力。在终身提示注入攻击流中,与最先进的防御措施相比,COPA将攻击成功率降低了高达6.3倍,平均降低了4.4倍。这些结果表明,持续偏好优化是防御LLMs免受自适应对手攻击的有效范式。
英文摘要
LLMs remain vulnerable to prompt injection attacks, where adversarial instructions embedded in user inputs or external content manipulate model behavior and bypass safeguards. Existing defenses are predominantly static, relying on fixed alignment objectives or attack-specific filtering mechanisms that require redesign as new attack strategies emerge. While recent lifelong alignment methods address shifting user preferences, they do not account for adaptive adversaries that continually evolve to exploit weaknesses in previously learned defenses. This limitation is particularly important in real-world deployments, where evolving attack distributions necessitate continual adaptation without sacrificing robustness to previously encountered threats. We present COPA, a continual preference optimization framework that treats prompt-injection defense as a lifelong learning problem. Instead of one-time alignment, COPA incrementally incorporates feedback from newly observed attacks via GRPO-based optimization and uses margin-weighted experience replay to retain defenses against prior attack classes. This enables continuous adaptation to emerging threats while mitigating catastrophic forgetting and preserving general-purpose model capabilities. Across lifelong prompt injection attack streams, COPA reduces attack success rate by up to 6.3x and 4.4x on average compared to state-of-the-art defenses. These results highlight continual preference optimization as an effective paradigm for defending LLMs against adaptive adversaries.