发表机构
The Pennsylvania State University; Purdue University(宾夕法尼亚州立大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对前沿LLM提示注入红队测试的冷启动问题,提出基于课程强化学习的方法,通过渐进训练攻击者LLM,在AgentDyn上实现高达93.8%的攻击成功率,并展现跨目标迁移能力。
AI 中文摘要
提示注入是LLM及基于LLM的应用(如智能体)面临的主要安全风险。针对提示注入的最先进红队测试方法利用强化学习(RL)训练攻击者LLM生成有效的注入提示。然而,当针对GPT-6-Luna等前沿LLM时,主要挑战是冷启动问题:攻击者LLM的每次攻击尝试均失败,从而获得零奖励,无法为学习提供信号。在本工作中,我们提出一种基于课程学习的方法来解决冷启动问题。具体而言,我们建议训练攻击者LLM对抗一系列鲁棒性逐渐增强的目标LLM,每个阶段从上一阶段获得的攻击者LLM进行热启动。然而,仅针对弱目标(如GPT-4o-mini)进行训练可能不足以使攻击者LLM获得针对前沿LLM(如GPT-5.6-Terra)的有用学习信号。相反,我们发现课程的设计至关重要:在每个阶段之后,攻击者LLM需要针对下一个目标LLM部分成功,以便其能够从成功尝试中学习以攻击新目标。我们的广泛评估表明,我们的方法能够有效对前沿LLM进行红队测试,在AgentDyn上对GPT-5.6-Luna和GPT-5.6-Terra的攻击成功率(ASR@10)分别达到93.8%和45.0%,而最先进的RL方法如RL-Hammer和PISmith在相同设置下ASR为0%。此外,我们发现攻击者LLM具有跨目标迁移能力,例如,训练用于击败一个强LLM(GPT-5.6-Terra)的攻击者LLM也能成功攻击六个其他前沿LLM(如GPT-6-Luna),而这些LLM从未在训练中出现。我们的代码可在\this https URL\this处获取。
英文摘要
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8\% and 45.0\% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0\% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at https://github.com/albert-y1n/PIForge.
Comments19 pages, 1 figure