JailMeter:一种基于证据的大语言模型越狱攻击评估框架
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Renmin University of China(中国人民大学)
- Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大语言模型越狱攻击评估标准和方法不一致的问题,提出基于证据评估框架JailMeter,受信息瓶颈理论启发用双反馈优化过滤噪声,在JailMeter-Eva上评估准确率达97.27%,还提炼出JailMeter\textsubscript{SLM}降低计算成本。
AI中文摘要:
当前针对大语言模型的越狱攻击评估存在评估标准和方法不一致的问题,导致攻击成功率估计不可靠。我们提出了JailMeter,一个基于证据的评估框架,旨在更准确地衡量越狱有效性。受信息瓶颈理论启发,JailMeter应用双反馈优化从模型响应中过滤越狱噪声,同时保留与原始恶意问题相关的内容。该过程产生简洁证据用于严格评估,只有当响应捕捉到恶意意图并给出完整答案时攻击才被验证。我们在包含330个人工标注、未被拒绝的越狱实例的JailMeter-Eva上评估JailMeter,其准确率达到97.27%,显著优于现有评估方法。为支持大规模评估,我们进一步将JailMeter提炼为一个小语言模型JailMeter\textsubscript{SLM},它在显著降低计算成本的情况下保持了相当的可靠性。代码和数据集可通过此https URL获取。
英文摘要:
The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness. Inspired by the Information Bottleneck theory, JailMeter applies dual-feedback optimization to filter jailbreak noise from model responses while preserving content relevant to the original malicious question. This process produces concise evidence for a rigorous assessment under which an attack is validated only when the response captures the malicious intent and delivers a complete answer, thereby signaling a substantive bypass of model safety alignment. We evaluate JailMeter on JailMeter-Eva, a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. JailMeter achieves an accuracy of 97.27%, substantially outperforming existing evaluation methods. To support large-scale evaluation, we further distill JailMeter into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs. Code and dataset are available at https://github.com/Magi2B0y/JailMeter.