arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08471cs.AI

昨日之盾,今日之矛:生产环境中的自演进安全护栏

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

Cong Ming, Jingyi Chen, Bin Liu, Qi Chu, Tao Gong, Nenghai Yu, Yingfei Xiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对静态LLM安全护栏无法应对新威胁的问题,提出自演进安全护栏SESG多智能体系统,可快速适配新威胁,性能优于静态护栏及自适应基线,已应用于深信服护栏并取得良好效果。

中文摘要 AI 辅助

已部署的大语言模型(LLM)安全护栏大多是静态的:仅训练一次并在发布时冻结,而新的越狱技术和此前未处理的有害类别会在数日内出现,导致防御始终落后一步。我们提出SESG(Self-Evolving Safety Guardrails,自演进安全护栏),一个运行在生产环境中的多智能体系统。SESG监控已部署护栏背后的实时流量,识别两类故障:形式新颖的越狱和内容新颖的有害类别。故障确认后,生成智能体会合成针对该故障的配对训练数据;验证智能体会重新平衡批次,使其偏向已部署模型出错的方向,让模型自身的错误引导其训练集;路由智能体会将训练动作与诊断出的缺口匹配,并返回下一版本投入生产。经过六轮实时演进(V0至V6),一个17亿参数的护栏能在16至24小时内适应新威胁,仅需约2小时的人工投入,而其替代的人工流程需40至90小时。在6种新出现的威胁上,它的性能优于0.6B至9B参数的静态护栏及自适应基线,同时保持了通用筛查能力。自2026年4月起,SESG已成为深信服(Sangfor)护栏的主要更新流水线,在两个月内自主关闭了15种新威胁场景中的14种。我们在该httpsURL发布了针对6种新威胁的9个测试集。警告:本文包含可能有害或冒犯性的示例。

英文摘要

Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed harmful categories emerge within days, leaving the defense perpetually a step behind. We present SESG (Self-Evolving Safety Guardrails), a multi-agent system running in production. SESG monitors the live traffic behind a deployed guardrail and surfaces two classes of failure: jailbreaks novel in form and harmful categories novel in content. Once a failure is confirmed, a generation agent synthesizes paired training data targeted at it; a validation agent rebalances the batch toward the direction in which the deployed model errs, so that the model's own mistakes steer its training set; and a routing agent matches the training action to the diagnosed gap and returns the next version to production. Over six rounds of live evolution (V0 to V6), a 1.7B guardrail adapts to a new threat in 16-24 hours, with about 2 hours of human effort, versus the 40-90 hours of the manual process it replaces. On six emerging threats, it outperforms static guardrails from 0.6B to 9B and an adaptive baseline while preserving its general screening competence. Since April 2026, SESG has been the primary update pipeline of Sangfor's guardrail, autonomously closing 14 of 15 new threat scenarios in two months. We release 9 test sets for the 6 new threats at https://github.com/Trams1017/SESG. Warning: This paper contains examples that may be harmful or offensive.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Shenzhen University(深圳大学)
  • Sangfor Technologies(深信服科技)

机构由 AI 辅助整理,请以论文原文为准。

↑