发表机构
University of Illinois Urbana-Champaign; Capital One(伊利诺伊大学厄巴纳-香槟分校; 第一资本金融公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有LLM越狱评估依赖语言合理性而非正确性的局限,提出SEAV框架,可降低假阳性率并重新归类大量先前标记为成功的越狱,显著改变测得的鲁棒性。
AI 中文摘要
越狱鲁棒性已成为大语言模型(LLM)安全评估的核心,但主流方法主要依赖拒绝行为、语义相似性和意图匹配启发式,强调语言合理性而非正确性。我们发现现有评估的关键局限:许多越狱意图依赖指令有效性而非认知事实性,使得看似真实的响应即便在事实或程序上不正确也被标记为成功。为解决此问题,我们提出序列认知与动作级验证(SEAV),这是一种以验证为核心的越狱评估框架,将响应分解为有序步骤,同时评估有效性与正确性。SEAV结合用于语义解释的LLM作为评判者机制,以及使用外部知识源的检索基础验证,评估生成内容是否事实正确、结构一致,且具备推进有害目标的操作能力。实验表明,SEAV在SD-A(精心整理的战略欺骗诊断工具)上将假阳性率较最强基线降低14.9个百分点,并在四个公共基准中的三个上,将22.1%至51.0%的先前标记为成功的抽样重新归类为无效。这些结果共同表明,强制正确性会显著改变测得的鲁棒性:许多先前标记为成功的越狱被重新归类为无效,且结果在测试的搜索后端和评估模型中保持稳定。代码和数据可在此https URL获取。
英文摘要
Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness. We identify a key limitation in existing evaluations: many jailbreak intents depend on instructional validity rather than epistemic factuality, allowing realistic-looking responses to be labeled successful despite being factually or procedurally incorrect. To address this gap, we propose Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness. SEAV combines LLM-as-a-judge mechanisms for semantic interpretation with retrieval-grounded verification using external knowledge sources, assessing whether generated content is factually correct, structurally consistent, and operationally capable of advancing harmful objectives. Empirically, SEAV cuts the false-positive rate on SD-A (a curated strategic-dishonesty diagnostic) by 14.9\,pp vs. the strongest baseline, and reclassifies 22.1\%--51.0\% of sampled prior-labeled successes as invalid across three of four public benchmarks. Together, these results show that enforcing correctness substantially reshapes measured robustness: many previously labeled jailbreak successes are reclassified as invalid, and results are stable across the tested search backends and evaluator models. Code and data are available at https://github.com/Ardor-Wu/SEAV.
CommentsTo appear on EMNLP 2026 main