发表机构
East China Normal University; Shanghai Innovation Institute(华东师范大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SRD-GUARD通过语义重写和联合多模型评分暴露隐藏意图,实现黑盒越狱防御,在多个基准上取得更优的DSR-ORR权衡。
AI 中文摘要
大型语言模型(LLMs)越来越多地部署在安全关键应用中,然而越狱攻击可以通过角色扮演、虚构场景或看似良性的动机来隐藏有害意图。现有的推理时防御可能无法识别伪装攻击,或过度拒绝合法请求。我们提出了SRD-GUARD,一个无需参数、黑盒的防御框架,通过语义重写和基于共识的风险评估来暴露隐藏意图。给定输入提示,SRD-GUARD生成五个语义相关的重写版本,这些版本保留底层目标,同时去除不必要的上下文包装。原始提示和重写版本由多个独立的基于LLM的安全评分器在连续风险尺度上联合评估。决策模块结合绝对风险阈值和原始与重写提示之间的相对风险变化,自适应地拦截、保留或警告请求。我们在Llama-3-8B-Uncensored和DeepSeek-V4-Flash上,使用AdvBench和OR-Bench-Hard,评估了SRD-GUARD对抗UNIATTACK、CIPHER和DeepInception的性能。SRD-GUARD实现了平均DSR分别为91.44%和100%,ORR分别为8.00%和12.00%。与评估的基线相比,它提供了更有利的DSR-ORR权衡。消融研究表明,重写暴露了隐藏的有害意图,联合评分提高了对单个评估器行为的鲁棒性,风险自适应决策能够选择性地处理模糊输入。这些结果表明,语义意图暴露、基于共识的风险评估和相对风险感知路由为黑盒越狱防御提供了一种有效且选择性的方法。该工件可在以下https URL获取。
英文摘要
Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inference-time defenses may miss disguised attacks or excessively refuse legitimate requests. We propose SRD-GUARD, a parameter-free, black-box defense framework that exposes concealed intent through semantic rewriting and consensus-based risk assessment. Given an input prompt, SRD-GUARD generates five semantically related rewrites that preserve the underlying objective while removing unnecessary contextual packaging. The original prompt and rewrites are jointly evaluated by multiple independent LLM-based safety scorers on a continuous risk scale. A decision module combines absolute risk thresholds with relative risk changes between the original and rewritten prompts to adaptively intercept, preserve, or warn on requests. We evaluate SRD-GUARD against UNIATTACK, CIPHER, and DeepInception on Llama-3-8B-Uncensored and DeepSeek-V4-Flash using AdvBench and OR-Bench-Hard. SRD-GUARD achieves average DSRs of 91.44% and 100%, with ORRs of 8.00% and 12.00%, respectively. Compared with evaluated baselines, it provides a more favorable DSR--ORR trade-off. Ablation studies show that rewriting exposes concealed harmful intent, joint scoring improves robustness to individual evaluator behavior, and risk-adaptive decision making enables selective handling of ambiguous inputs. These results demonstrate that semantic intent exposure, consensus-based risk assessment, and relative-risk-aware routing provide an effective and selective approach to black-box jailbreak defense. The artifact is available at https://anonymous.4open.science/status/CICD-Guard-D648.