arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17829cs.CRcs.AI

模型的“信号泄露”:用行为度量器测量上下文泄露攻击信号

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 LeakGauge,通过附加后缀度量泄露行为并映射预填充 token 概率为攻击风险分数,在 11 个 LLMs 上对未见过的攻击达到 0.944-0.996 的 AUROC,可实现低延迟的输入检测器。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越依赖外部上下文,例如预定义的系统提示或检索到的文档,以提升生成质量。然而,在处理用户查询的同时处理这些上下文会产生攻击面:对抗性输入可诱导模型泄露这些上下文。先前的探测研究表明,与泄露相关的信号出现在隐藏状态中,但提取这些状态会带来额外的部署挑战。本文探讨这种内部信号是否在解码前留下更易获取的“信号”。我们提出 LeakGauge,它通过附加后缀来探测响应,该后缀可度量泄露行为,并将其预填充 token 概率映射为攻击风险分数。虽然直接度量使用机密内容的初始 token,但我们发现,一种与内容无关的、将泄露行为表述出来的度量方法能产生更稳健的信号。在包括 GLM-5.2(753B)和 Kimi-K3(2.8T)在内的 11 个 LLMs 上,LeakGauge 在未见过的攻击上达到 0.944-0.996 的 AUROC 范围。当内容更改语言或攻击从逐字泄露变为语义泄露时,该信号保持稳定。通过激活导向干预,我们进一步表明风险分数对内部泄露相关方向敏感,将可观测信号与模型的内部表示关联起来。此外,LeakGauge 可实现一个仅需少于 0.5K 额外参数、增加 10.34ms 延迟的输入检测器。代码:this https URL。

英文摘要

LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states poses additional deployment challenges. In this paper, we explore whether this internal signal leaves a more accessible ``tell'' before decoding. We propose LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack-risk score. While a direct gauge uses the initial tokens of confidential content, we find that a content-agnostic one that verbalizes leakage behavior yields more robust signals. Across 11 LLMs, including GLM-5.2 (753B) and Kimi-K3 (2.8T), LeakGauge reaches an AUROC range of 0.944--0.996 on unseen attacks. The signal remains stable when the content changes language or the attack shifts from verbatim to semantic disclosure. By activation-steering interventions, we further show that the risk score is sensitive to an internal leakage-related direction, relating the observable signal to the model's internal representation. In addition, LeakGauge enables an input detector with fewer than 0.5K extra parameters and added latency of 10.34 ms. Code: \href{https://github.com/yeasen-z/LeakGauge}.

发表机构

  • Tsinghua University(清华大学)
  • Ant International(蚂蚁国际)
  • Nanyang Technological University(南洋理工大学)
  • SiliconProspect AI(硅景人工智能)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑