加密语言模型解码的静态自举放置
Static Bootstrap Placement for Encrypted Language Model Decoding
浏览论文内容
中文总结 AI 辅助
针对同态加密下语言模型解码的服务器端循环问题,提出AR-HE系统,通过静态自举放置规则和加密缓存优化,将GPT-2小规模令牌生成成本从4715秒降至544秒。
中文摘要 AI 辅助
语言模型越来越多地服务于携带私有数据的提示词,而同态加密下的安全推理允许客户端在不泄露提示词的情况下外包计算。现有的安全推理系统在加密状态下执行前向传播时不消耗令牌,使用它们生成文本需要在每个生成的令牌处进行客户端往返。相反,将循环保留在服务器端需要在加密状态下选择并消耗令牌,并为随上下文增长的循环体放置自举。我们构建了AR-HE,它在服务器上运行整个循环,在加密状态下选择每个令牌,检索其嵌入,并将其写回加密状态。客户端发送一个提示词,然后保持离线直到输出。一条规则无需搜索即可放置运行中的每个自举,因此令牌的自举成本是上下文长度的函数,该函数在运行开始前已知。该调度跳过其结果无法到达输出的工作,打包共享操作数的自举,并将过去位置的键和值保存在加密缓存中。应用所有优化后,在一台NVIDIA H100上生成一个GPT-2小规模令牌需要544秒,而未优化时为4715秒。其之前的提示词步骤需要4630秒。仅缓存一项就将生成步骤从11751次自举减少到1072次。该公式预测了我们测量的每一步,包括它并非从中推导出的模型的步骤。
英文摘要
Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generating text with them requires a client round trip at every generated token. Keeping the loop on the server instead requires selecting and consuming a token under encryption, and placing bootstraps for a loop body that grows with the context. We build AR-HE, which runs the whole loop on the server, selects each token under encryption, retrieves its embedding, and writes it back into the encrypted state. The client sends one prompt and remains offline until the output. One rule places every bootstrap in the run, without search, so the bootstrap cost of a token is a formula in the context length that is known before the run starts. The schedule skips work whose result cannot reach the output, packs bootstraps that share an operand, and keeps the keys and values of past positions in an encrypted cache. With every optimization applied, generating a GPT-2 small token costs 544 seconds on one NVIDIA H100, down from 4715 seconds without optimization. The prompt step before it costs 4630 seconds. The cache alone takes a generated step from 11751 bootstraps to 1072. The formula predicts every step we measured, including steps of a model it was not derived from.
发表机构
- Koç University(科奇大学)
机构由 AI 辅助整理,请以论文原文为准。