发表机构
MLCommons; University of Warwick; Qualcomm; AVERI; The Alan Turing Institute; Google DeepMind; New York University; Argonne National Laboratory; University of Illinois Urbana-Champaign; Monark Health; NVIDIA; Microsoft; Rama Labs; Think Evolve Labs; New Jersey Institute of Technology; Clarkson University; Polytechnique Montréal(MLCommons; 华威大学; 高通公司; AVERI; 艾伦·图灵研究所; 谷歌DeepMind; 纽约大学; 阿贡国家实验室; 伊利诺伊大学厄巴纳-香槟分校; Monark Health; 英伟达; 微软; Rama Labs; Think Evolve Labs; 新泽西理工学院; 克拉克森大学; 蒙特利尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MLCommons 越狱基准 v1.0 提出了一种评估大语言模型对单轮文本越狱攻击鲁棒性的端到端方法,通过基线与对抗性评估及韧性差距量化安全性能变化,为比较性越狱评估建立了可复现基础。
AI 中文摘要
现代人工智能系统被设计用于拒绝危险请求。越狱是一种精心设计的提示,旨在绕过这些安全防护措施,并引出系统通常拒绝提供的输出。MLCommons 越狱基准 v1.0 提供了一种端到端的方法论,用于评估大型语言模型对单轮、基于文本的越狱攻击的鲁棒性。它在一个基准测试流程中整合了基于标准的系统和攻击选择、配对的基线与对抗性评估、人工标注、自动化评估器校准、评分、分级以及风险校准披露。该基准使用 264 个种子提示评估了八个开放权重系统,这些提示涵盖十一个危害类别,并包含从 MLCommons 越狱分类法中提取的代表性攻击。响应使用 AILuminate 评估标准 v1.4 进行评估,鲁棒性通过韧性差距来衡量:即基线与对抗性条件下安全性能的变化。在所有评估的系统和攻击中,不安全响应率从基线条件下的 11.08% 增加到越狱条件下的 18.65%,产生了 7.57% 的平均韧性差距。可访问的系统显示出更大的平均差距,而攻击有效性在不同攻击类别和危害之间差异显著。该基准还考察了评估器的可靠性以及测量误差的来源。除了报告结果之外,越狱基准 v1.0 还建立了一个可复现的方法论基础,用于比较性越狱评估,以及未来在系统、攻击、危害和评估方法上的扩展。
英文摘要
Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to single-turn, text-based jailbreak attacks. It combines criteria-driven system and attack selection, paired baseline and adversarial evaluation, human annotation, automated evaluator calibration, scoring, grading, and risk-calibrated disclosure within a single benchmarking pipeline. The benchmark evaluates eight open-weight systems using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy. Responses are assessed using the AILuminate Assessment Standard v1.4, and robustness is measured through the Resilience Gap: the change in safety performance between baseline and adversarial conditions. Across all evaluated systems and attacks, the unsafe-response rate increased from 11.08% under baseline conditions to 18.65% under jailbreak conditions, producing an average Resilience Gap of 7.57%. Accessible systems showed a larger mean gap, while attack effectiveness varied substantially across attack categories and hazards. The benchmark also examines evaluator reliability and sources of measurement error. Beyond reporting results, Jailbreak Benchmark v1.0 establishes a reproducible methodological foundation for comparative jailbreak evaluation and for future expansion across systems, attacks, hazards, and evaluation methods.