arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对黑盒大语言模型的实际秘密提取

Practical Secrets Extraction against Black-box LLMs

Shiqian Zhao, Siwei Jiang, Xinfeng Li, Runyi Hu, Yandan Zheng, Congyu Guo, Tianwei Zhang, Anh Tuan Luu

arXiv 2609.36941首次发表:更新:

发表机构

Nanyang Technological University; Beijing University of Posts and Telecommunications; Hong Kong Polytechnic University(南洋理工大学; 北京邮电大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种针对商业黑盒LLM的秘密提取框架,通过知识蒸馏和代理引导过滤,在仅输出访问下有效恢复记忆化凭据,提升提取效率与真实密钥率。

AI 中文摘要

大型语言模型(LLMs)日益为自主编码智能体(如Codex和Claude Code)提供动力,然而其训练语料库可能包含在公共仓库中暴露或从私有开发工件中收集的机密凭据,从而造成记忆化和随后泄露的风险。然而,现有的提取审计大多假设能够访问模型权重或令牌概率。在这项工作中,我们提出了一种针对商业、基于API的LLMs在仅输出访问条件下的黑盒秘密提取框架。该框架包括:(i)交叉验证的秘密知识蒸馏,它使用保持语义的提示变体、响应交叉验证和特定于提供商的格式过滤,将秘密相关行为蒸馏到本地白盒代理中;(ii)代理引导的秘密提取和候选过滤,它结合截断的top-p采样与本地令牌熵、N-gram频率分析和特定于提供商的结构先验。在受控的API密钥基准测试中,我们的框架在恢复有效性和真实密钥率方面优于代表性基线,同时降低了提取延迟。一项负责任的真实世界评估进一步从三个独立部署的黑盒LLM系统中恢复了被掩蔽的提供商特定凭据,涵盖OpenAI和Claude Code,表明在仅输出访问条件下可以暴露记忆化的秘密。

英文摘要

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering}, which combines truncated top-$p$ sampling with local token entropy, $N$-gram frequency profiling, and provider-specific structural priors. On controlled API-key benchmarks, our framework improves recovery effectiveness and real-key rates over representative baselines while reducing extraction latency. A responsible real-world evaluation further recovers masked provider-specific credentials from three independently deployed black-box LLM systems spanning OpenAI and Claude Code, showing that memorized secrets can be exposed under output-only access.

CommentsThis paper proposes a practical secret extraction method against black-box large language models

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑