arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPT-Red:基于大规模自博弈的自动化红队测试

GPT-Red: Automated Red Teaming via Self-Play at Scale

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen

arXiv 2607.26115首次发表:更新:

发表机构

OpenAI(OpenAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出GPT-Red,一种基于大规模自博弈的自动化红队测试智能体,可发现新型提示注入攻击、攻破过往模型且泛化能力强,用于提升LLM鲁棒性并形成自我改进飞轮。

AI 中文摘要

我们推出GPT-Red,这是一种自动化红队测试智能体,旨在发现针对前沿大语言模型(LLM)的新型提示注入攻击,其目标是评估并提升我们生产系统的鲁棒性。为此,我们利用它对GPT-5.6进行对抗训练,GPT-5.6是我们迄今为止在提示注入方面最鲁棒的模型。为构建GPT-Red,我们设计了一种可扩展的自博弈算法,该模型的任务是攻击同时训练的多样化防御智能体群体。我们在逼真的红队测试环境中,使用与我们部分最大规模的RL后训练运行相当的计算资源对该模型进行训练,这使其成为有记录以来规模最大的LLM安全训练运行。GPT-Red在红队测试中表现出色:它能可靠地攻破我们过往至GPT-5.5的模型,发现比人类红队测试人员更多的成功攻击,且能泛化到未见过的环境、防御模型和工具中。未来,我们预计随着每款新GPT模型鲁棒性的提升,它将为更强大的红队测试智能体提供更优质的学习信号,从而开启自我改进的飞轮效应。

英文摘要

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.

Comments28 pages.13 main pages and 13 main figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑