从专有大语言模型API窃取推理轨迹
Stealing Reasoning Traces from Proprietary LLM APIs
浏览论文内容
中文总结 AI 辅助
该研究发现专有LLM API的加密推理轨迹存在跨会话、用户及模型的兼容性漏洞,开发了可扩展解密越狱方法,能提取模型推理、窃取PII等,还提出了对应缓解措施。
中文摘要 AI 辅助
主流大语言模型提供商如今会隐藏其模型的逐步推理过程(即思维链),以保护知识产权并限制信息泄露。提供商不会将这些推理轨迹存储在服务器端,而是将其作为加密文本块返回给客户端,客户端会在后续每次请求中附带这些加密块。基于现有研究,我们发现了一个架构漏洞:这些加密块在提供商生态系统内的不同会话、用户和模型之间完全兼容且可互换。我们利用这种兼容性开发了一种可扩展的解密越狱方法:将给定模型的加密推理轨迹注入同一提供商旗下一个较弱、防护措施较少的模型中,迫使其逐字解码并以明文输出该轨迹,无需直接对功能更强的模型进行越狱。该漏洞可实现四种不同的攻击途径:第一,规避抗蒸馏机制,使攻击者能够提取专有模型的推理过程,我们在Anthropic、OpenAI和Google平台上均验证了这一点;第二,允许大规模提取私人数据,开发者常公开共享会话日志却未察觉加密块的内容,通过解码从公共仓库爬取的315320个推理块,我们恢复了367项个人身份信息(PII)工件和182个凭证;第三,会意外泄露推理过程中隐藏的危险信息,即便模型最终可见输出已安全拒绝恶意请求;第四,攻击者可利用此漏洞执行隐形提示注入,将恶意有效载荷完全嵌入加密块中,以毒害公共智能体部署。在负责任披露后,我们提出了具体的密码学及系统级缓解措施,以保护客户端推理过程安全。
英文摘要
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.
发表机构
- MATS Research(MATS研究院)
- ELLIS Institute Tübingen(图宾根ELLIS研究所)
- Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
- Tübingen AI Center(图宾根人工智能中心)
- University of Tübingen(图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。