越狱评估器的实证测量
An Empirical Measurement of Jailbreaking Evaluators
AI总结:
本研究系统比较六个越狱评估器,以人工判断为基准,发现JADES性能最佳,HarmBench和StrongReject表现良好,为评估器选择提供实证依据。
AI中文摘要:
专家对越狱响应的评估成本高昂且难以扩展,因此社区越来越依赖自动化评估器来判断攻击是否成功。然而,越狱研究通常独立验证其选定的评估器,在类似的评估工作上反复消耗资源,同时使得不同论文的结果难以比较。不同的评估器也编码了不同的越狱成功定义,这意味着报告的攻击强度和表面进展可能在很大程度上取决于使用哪个评估器。我们系统比较了近期越狱攻击与防御研究中反复出现的六个评估器:HarmBench、JailbreakBench、JailbreakRadar、StrongReject、JADES和JailMeter。据我们所知,此前没有研究在受控设置下对同一人工标注数据评估全部六个评估器。我们在JailbreakQR和JailMeter-Eva上以人工判断为参照对它们进行评估,并测量与人类的一致性、错误类型以及跨攻击族的一致性。对于需要通用大语言模型评判员的评估器,我们使用共享骨干来控制模型特定差异。我们发现JADES表现出最佳整体性能,而HarmBench和StrongReject也展现出良好性能。
英文摘要:
Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their chosen evaluator independently, repeatedly spending resources on similar evaluation efforts while making results across papers difficult to compare. Different evaluators also encode different definitions of jailbreak success, meaning that reported attack strength and apparent progress can depend substantially on which evaluator is used. We systematically compare six evaluators that recur in recent jailbreak attack and defense research: HarmBench, JailbreakBench, JailbreakRadar, StrongReject, JADES, and JailMeter. To our knowledge, no prior study has evaluated all six on the same human-labeled data under a controlled setup. We evaluate them on JailbreakQR and JailMeter-Eva, using human judgments as the reference, and measure agreement with humans, error types, and consistency across attack families. For evaluators that require a general-purpose LLM judge, we use a shared backbone to control for model-specific variation. We found that JADES exhibits the best overall performance, while HarmBench and StrongReject also demonstrate good performance.