SpliTEE:利用差分隐私GPU外包提升可信硬件上的LLM推理
SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy
浏览论文内容
中文总结 AI 辅助
SpliTEE通过差分隐私保护中间表示,将LLM推理分割到TEE和GPU,实现比CPU推理快近两倍且比加密方案更快的安全加速。
中文摘要 AI 辅助
提供给大型语言模型(LLM)的用户提示可能包含敏感或私人信息,这些信息可能被远程部署的模型滥用,例如在重新训练期间的无意记忆。保护用户提示的一种方法是在可信执行环境(TEE)内执行LLM,并保证服务提供商无法访问TEE内执行的计算或与TEE交换的信息。然而,当前的TEE主要基于CPU,并且比针对LLM推理优化的GPU慢得多。为了规避这一问题,Tramer和Boneh(2019)提出了Slalom,它将神经网络推理在TEE和不可信GPU之间进行分割,并对发送到GPU的中间输入进行加密。我们将这种分割推理架构扩展到LLM推理,并转而使用差分隐私来保护中间输入。我们通过展示提示重建攻击能够以近80%的准确率从这些表示中恢复提示,证明了掩蔽中间表示的必要性。我们的主要贡献是对关键LLM函数的全局敏感性分析,这界定了差分隐私噪声所需的规模。与加密不同,差分隐私避免了量化,使LLM能够保持在浮点域中。我们还推导了TEE中掩蔽和噪声消除引起的浮点误差的上界,该上界是隐私参数ε的函数。我们使用Intel TDX实现了我们的架构,并使用两个LLM进行了评估:Llama-3.2-3B和Qwen3-4B。我们的分割执行比TDX内完全基于CPU的推理快近两倍,并且比基于加密的Slalom快5-15秒,同时实现了更高的准确性。最后,我们证明,即使知道差分隐私机制,提示重建也无法恢复比无关提示中包含的更多信息。
英文摘要
User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a trusted execution environment (TEE). However, this results in slow inference times as current TEEs are significantly slower than GPUs for LLM inference. To circumvent this, Tramèr and Boneh (2019) proposed Slalom which splits neural network inference between a TEE and an untrusted GPU. They encrypt inputs to computations outsourced to the GPU. In this paper, we extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy (DP). We first demonstrate that masking intermediate representations is necessary by showing an 80% accuracy on a prompt-reconstruction attack from these representations. Our main contribution is a global sensitivity analysis of key functions in LLM inference, which bounds the required scale of DP noise. Unlike encryption, DP avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on the floating-point error from masking and subsequent noise cancellation as a function of the privacy parameter epsilon, keeping the same quality of the LLM response. We implement our architecture using the Intel TDX TEE and two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully TDX-based inference. Moreover, it is at most 43% faster than Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the DP mechanism, cannot recover more information than is contained in an unrelated prompt.
发表机构
- Macquarie University(麦考瑞大学)
机构由 AI 辅助整理,请以论文原文为准。