发表机构
Virtue AI; Carnegie Mellon University; Northeastern University; University of Chicago; University of Illinois, Urbana-Champaign(美德人工智能公司; 卡内基梅隆大学; 东北大学; 芝加哥大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基础模型安全基准测试随模型和法规发展而不足的问题,提出AIR-BENCH Live,通过自动化更新管道和多智能体算法更新提示,扩展了基准测试风险数量,评估模型发现安全范围广等结果,使其能随领域发展。
AI 中文摘要
基础模型安全基准测试反映了其发布时的人工智能风险,但随着模型改进和新法规出台,其风险分类法变得不全面,攻击提示也变得无效。我们提出了AIR-BENCH Live,它是AIR-BENCH 2024的自我演进继任者。自动化更新管道监测政府法规,并根据当前的四级风险分类法对新政策进行分类,将其与现有类别匹配或提出新的细分类别。然后,一种多智能体、基于角色的提示生成算法在最少人工审核的情况下生成逼真的多语言提示,同时为现代越狱技术留下改进空间。该算法用于更新旧提示并为新类别生成提示。在当前版本中,管道已将基准测试从314个细分风险扩展到335个,21个新类别来自七个司法管辖区的31条真正新颖的政策条款。评估14个最新模型时,我们发现安全范围很广(根据模型自身行为判断,模型之间从0.17到0.89),现代化提示平均比2024年的设置难0.06分,最大降幅集中在最合规的模型中,并且大多数模型在非英语提示上的安全性略低。通过不断吸收新法规并重新生成提示,AIR-BENCH Live旨在与快速发展的领域同步发展。
英文摘要
Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.
Comments11 pages, 7 figures, 3 tables