去令牌化泄露:从缓存踪迹重建本地大语言模型输出
Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces
- Ben Gurion University(本·古里安大学)
- Microsoft Security(微软安全部门)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
通过缓存侧信道攻击去令牌化器,利用Flush+Reload和Prime+Probe重建本地大语言模型输出,并在多平台验证,影响广泛。
AI中文摘要:
我们提出了一种新的攻击方法,通过观察去令牌化过程中的CPU缓存活动,重建由本地托管的大语言模型生成的文本。与以往依赖特定部署假设(如共享数据内存、CPU卸载或混合专家架构)的攻击不同,我们的方法针对去令牌化器,这是默认大语言模型推理流程中使用的一个组件。为了获得清晰的信号,我们在共享分词器代码上使用Flush+Reload来检测解码发生的时间,这使我们能够在正确的时刻执行Prime+Probe,并隔离与令牌相关的缓存活动。然后,我们应用聚类和语言模型流水线,从嘈杂的缓存观察中恢复文本。我们在多个数据集、硬件平台、推理框架和模型系列上评估了该攻击,并表明它可以从真实的本地大语言模型部署(包括智能体系统)中恢复语义准确的输出。这一漏洞尤其重要,因为最广泛使用的分词器实现容易受到该攻击,并且嵌入在许多流行的本地大语言模型产品和智能体框架中,包括我们演示的OpenClaw等系统,从而大幅拓宽了实际攻击面。
英文摘要:
We present a new attack that reconstructs the text generated by locally hosted LLMs by observing CPU cache activity during detokenization. Unlike prior attacks that rely on deployment-specific assumptions, such as shared data memory, CPU offloading, or Mixture-of-Experts architectures, our approach targets the detokenizer, a component used in default LLM inference pipelines. To obtain clean signals, we use Flush+Reload on shared tokenizer code to detect when decoding occurs, which lets us perform Prime+Probe at the right moment and isolate token-dependent cache activity. We then apply a clustering-and-language-model pipeline to recover text from noisy cache observations. We evaluate the attack across multiple datasets, hardware platforms, inference frameworks, and model families, and show that it can recover semantically accurate outputs from real-world local LLM deployments, including agentic systems. This vulnerability is particularly significant because the most widely used tokenizer implementations are susceptible to the attack and are embedded in many popular local LLM products and agent frameworks, including systems such as OpenClaw (which we demonstrate), substantially broadening the practical attack surface.