ReSI:面向抗性与韧性AI的递归安全改进
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
AI总结:
该研究提出ReSI递归安全改进框架,通过多轮自动化评估更新模型,在保留通用能力的同时提升安全性能,降低X-Teaming攻击成功率,实现了抗性与韧性AI的安全对齐。
AI中文摘要:
递归自我改进,即AI系统参与提升自身能力,正从理论展望走向实践,为安全对齐带来了挑战与机遇。模型通过频繁更新实现演进,其安全对齐需要持续适配每个新检查点。同时,随着不断演进的红队方法暴露出新漏洞,每个检查点的安全改进需要缓解已暴露的漏洞,并泛化到尚未显现的风险。遵循R²AI,我们将这些目标称为对已知威胁的抗性和对未预见风险的韧性。递归自我改进进而为实现这两个目标提供了一种方法:安全对齐同样可以通过连续多轮评估与更新推进。因此,我们提出ReSI,一种通过自动化研究实现该方法的递归安全改进框架。在每一轮中,ReSI应用多样化的红队方法识别当前目标模型的漏洞,开发训练方案,并在通过能力保留帕累托门限的更新中,将安全增益最大的更新提升为下一个目标模型。在四个密集型模型和混合专家模型上,ReSI在分布内和分布外安全基准上与评估的前沿模型相当或更优,且在几乎所有安全评估中优于对齐基线,同时在很大程度上保留了通用能力。特别地,ReSI将四个模型的平均X-Teaming攻击成功率从86.01%降至31.45%,远低于前沿模型GPT-5.6-Luna的领先结果56.69%,表明其对训练期间未见过的攻击具有更强的韧性。这些发现支持递归安全改进是通向抗性与韧性AI的可行路径。
英文摘要:
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R$^2$AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.