arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新激活对齐:通过意图感知的输入-输出匹配防御大语言模型的越狱攻击

Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching

Luoyu Chen, Weiqi Wang, Chenhan Zhang, Zhiyi Tian, Shui Yu

arXiv 2610.04470首次发表:更新:

发表机构

University of Technology Sydney; Southeast University(悉尼科技大学; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SENTINEL,一种基于意图提取的生成时防御方法,利用输入-输出语义一致性检测并阻止越狱攻击,在HarmBench上将成功率降至约5%且过度拒绝率低。

AI 中文摘要

大型语言模型(LLMs)仍然容易受到越狱攻击,这些攻击将有害意图隐藏于复杂的对抗性提示中。现有的防御主要依赖于输入扰动或有害输出抑制,但很少对恶意意图所在位置进行建模,导致防护脆弱且过度拒绝现象严重。我们提出SENTINEL,一种即插即用、生成时进行的越狱防御方法,将缓解问题重新定义为意图提取问题。我们的关键见解是,经过指令微调的LLM表现出强烈的输入-输出语义一致性:无论越狱复杂性如何,生成的输出往往与攻击者的真实意图保持一致。SENTINEL利用这一特性,通过匹配语义对齐的输入-输出区域来提取揭示意图的子序列,使用拒绝方向投影对这些子序列进行评分以估计有害程度,并在必要时停止生成。在HarmBench上对多个LLM进行的实验表明,SENTINEL将越狱成功率降低至接近5%,同时保持较低的过度拒绝率。我们进一步展示了对自适应攻击的鲁棒性,并提供了一种机制性解释:SENTINEL将越狱特征从对齐盲区重新分配到对齐区域。

英文摘要

Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive over-refusal. We propose SENTINEL, a plug-and-play, generation-time jailbreak defense that reframes mitigation as an intent extraction problem. Our key insight is that instruction-tuned LLMs exhibit strong input--output semantic consistency: regardless of jailbreak complexity, generated outputs tend to align with the attacker's true intent. SENTINEL exploits this property by matching semantically aligned input--output regions to extract intention-revealing subsequences, scores these subsequences using refusal-direction projections to estimate harmfulness, and halts generation when necessary. Experiments on HarmBench across multiple LLMs show that SENTINEL reduces jailbreak success rates to close to 5\% while maintaining low over-refusal. We further demonstrate robustness to adaptive attacks and provide a mechanistic interpretation: SENTINEL re-distributes jailbreak features from alignment blind spots to aligned regions.

Commentsemnlp2026 main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑