arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20723cs.CRcs.LG

泄漏语言模型:通过逐令牌计时窃取架构和推理优化

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram

首次发表
浏览论文内容

中文总结 AI 辅助

研究通过LeakyLMs攻击从生产语言模型中窃取信息,利用令牌生成时间推断关键细节。核心方法是构建令牌生成时间模型,可针对推理优化和部署策略及恢复架构属性。主要贡献是能有效窃取模型、架构和部署信息,实验效果良好。

中文摘要 AI 辅助

这项工作提出了LeakyLMs,这是一组从生产语言模型中泄漏专有模型、架构和部署信息的攻击方法。LeakyLMs首次证明,即使通过远程API进行交互,也可以仅使用令牌生成时间来推断关键模型和部署细节。LeakyLMs引入了两种核心攻击。第一种攻击针对推理优化和部署策略,比如能检测提供商是否使用推测性解码,并识别流水线中草稿模型的上下文长度。第二种攻击可恢复关键架构属性。为实现这些,LeakyLMs在现代NVIDIA GPU上构建了详细准确的令牌生成时间模型,利用该模型在架构空间中搜索。在对Llama模型的实验中,接近正确的架构配置在超过90%的情况下出现在前10个猜测中。

英文摘要

This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models. LeakyLMs is the first to demonstrate that key model and deployment details can be inferred using only token generation timing, even when interacting through remote APIs. LeakyLMs introduces two core attacks. The first attack targets inference optimizations and deployment strategies. For example, our attack detects whether a provider uses speculative decoding, a widely deployed inference-time optimization, and further identifies the context length of the draft model used in the pipeline. Our measurements show that Google Gemini Flash 2.5 uses speculative decoding with a draft context window of approximately 128K tokens. The second attack recovers key architectural properties, including the number of transformer layers, hidden dimension size, and number of attention heads. To achieve this, LeakyLMs builds a detailed and accurate model of token-generation timing on modern NVIDIA GPUs, characterizing how latency scales with model configuration and hardware parameters. The attack then performs a search over the architecture space using this timing model. In experiments with Llama models, the near-correct architectural configuration appears in the top-10 guesses more than 90% of the time.

发表机构

  • Purdue University(普渡大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑