arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07532cs.CRcs.AIcs.CL

通过暗知识中的模型无关潜在安全信号保护大语言模型

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

提出LADE方法,利用首token概率分布中的暗知识提取模型无关的潜在安全信号,通过kNN判别防御越狱攻击,在多种LLM上降低攻击成功率并保持安全-效用平衡。

中文摘要 AI 辅助

大语言模型(LLMs)发展迅速,引发了对其安全性的日益关注。近期工作提出了检测和防御攻击的方法,包括利用模型隐藏状态的解码阶段防御。然而,现有的解码阶段防御存在两个局限。首先,它们在安全性与过度拒答之间引入了权衡,加强安全性会降低模型在良性查询上的有用性。其次,许多此类方法依赖内部隐藏状态,因此受限于特定架构,导致大量开销且跨模型泛化能力有限。为解决这些局限,我们提出了LADE(用于防御的潜在安全信号),它利用通过对比有害与良性查询从暗知识(即输出概率分布中超出其argmax所携带的信息)中提取的潜在安全信号,这些信号位于首个token的输出概率分布中。我们的关键洞见是,在表面拒答token之外,首token分布中的暗知识包含潜在安全信号,定义为在有害与良性查询之间概率差异显著的token。我们证明这些信号在不同LLM间一致对齐,形成一种源于安全对齐的模型无关方向。LADE由三个组件组成:(1)从暗知识中提取潜在安全信号,从首token概率分布中选择top-k安全判别性token;(2)分词器映射,将这些token映射到不同分词器以实现模型无关应用;(3)基于kNN的判别,通过映射token上的k-最近邻搜索对查询进行分类。在多种LLM和基准测试中,LADE对广泛的越狱攻击具有鲁棒性,降低了攻击成功率,同时保持了有竞争力的安全-效用权衡。

英文摘要

LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We show that these signals consistently align across LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and benchmarks, LADE is robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.

发表机构

  • UNC Chapel Hill(北卡罗来纳大学教堂山分校)
  • Kyung Hee University(庆熙大学)
  • AIM Intelligence(AIM智能公司)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑