AI 中文总结
本研究推出AI安全防护最低标准V1.0,测试Claude Fable5等四款模型的越狱鲁棒性,发现模型间防御能力差异大,部分模型未出现通用越狱,可通过现有技术缩小防御差距。
AI 中文摘要
前沿AI模型开发者越来越依赖分层防护措施来防止灾难性滥用,但几乎没有公开证据表明这些防护措施能提供多少保护,或在不同开发者之间的一致性如何。我们推出了AI安全防护最低标准V1.0:包含67种易获取的静态越狱技术的分类法、将这些技术组合成超大型攻击空间的方法,以及旗舰模型针对该空间样本的基准测试。我们在两个互补数据集(共360个攻击者目标,涵盖化学、生物、放射/核与爆炸(CBRNE)威胁以及进攻性网络领域)上评估Claude Fable 5、GPT-5.6 Sol、Gemini 3.1 Pro和Grok 4.5,采用三阶段漏斗法识别通用越狱:能在某领域超过75%目标上引出符合操作要求响应的单一提示模板。我们还引入了直接模拟攻击者成本的越狱成本指标,对未发现通用越狱的情况采用右删失下界。各模型的鲁棒性差异极大:破解这些模型的成本相差超百倍。对我们的技术库进行随机搜索,发现针对Grok 4.5的63个通用越狱和针对Gemini 3.1 Pro的18个通用越狱,每个越狱的平均成本约为58美元和278美元;专家引导的组合将成本提升至385美元和231美元。无论采用哪种策略,Claude Fable 5和GPT-5.6 Sol均未产生任何通用越狱。由于满足最低标准仅需其他地方已公开描述和部署的防御措施,这些差距可通过现有技术缩小。我们建议采用结合推理、激活和输入/输出监控的纵深防御。结果维护于该http URL。
英文摘要
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for security: meeting it does not guarantee a secure model, but failing to meet it guarantees a lack of state-of-the-art security. Version 1.0 covers severe misuse requests across chemical, biological, radiological, nuclear, and explosive (CBRNE) threats and offensive cybersecurity. In this report, we tested four leading models for universal jailbreaks in the context of this minimal standard, and found more than a hundredfold difference in security. Claude Fable 5 and GPT-5.6 Sol held against every attack we ran, with no universal jailbreak found; we estimate they would likely cost more than \$14,200 to jailbreak, if it is possible with this methodology at all. Meanwhile, we found hundreds of universal jailbreaks for Grok 4.5 and Gemini 3.1 Pro; each broke for under \$300, with universal jailbreaks in Grok's weakest domain, cybersecurity, accessible for as little as \$24. The gap is fixable: every weakness we found belongs to a known class of attack that already has a defense deployed in production models. The leaderboard will be updated on a rolling basis as new models are released, and the evaluation methodology and Minimal Standard will be periodically revised to take into account the latest capabilities and the state-of-the-art in safeguards. The leaderboard is available at leaderboard.far.ai.