arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于隐私保护 Llama 3 8B 推理的开源端到端 FHE 实现

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Yuhang Fan, Yusi Chen, Kanyu Ye, Zhuoran Ji

arXiv 2609.12378首次发表:更新:

AI 中文总结

针对云 LLM 推理的隐私风险,提出 FHE 推理系统 Odin,通过协同设计密文打包与模型执行,在 Llama-3-8B 上实现 4.51 倍加速,为首个开源端到端 GPU CKKS 实现。

AI 中文摘要

云 LLM 服务通常要求用户将提示发送给模型提供商,从而产生隐私风险。全同态加密(FHE)允许服务器在不解密输入的情况下执行推理,但将数据表示为密文会增加存储和计算开销。在基于 CKKS 的 LLM 推理中,打包方案将逻辑张量映射到密文和槽位。因此,它决定了密文数量和线性层的同态成本,并限制了数据在线性层、注意力机制和非线性计算之间的传递方式。随着模型和序列的增长,低效的布局会累积编码、计算和布局转换的开销。我们提出了 Odin,一个为 Llama 协同设计密文打包和模型执行的 FHE 推理系统。从以权重编码为瓶颈的 THOR 风格基线出发,Odin 使用特征主序的跨层布局来统一残差连接和层接口,并为线性投影和注意力构建瞬时的算子内布局。这减少了宽投影中权重的冗余明文编码。在注意力机制内部,QK^T 产生的分数可直接供 Softmax 使用,而 PV 消耗由此产生的概率,避免了中间重新打包。对于非线性操作,我们使用 minimax 多项式逼近,并采用输入范围控制和基于模型质量指导的联合误差分配,从而降低多项式次数和乘法深度。据我们所知,Odin 是首个开源的端到端 GPU CKKS Llama-3 实现。使用 Llama-3-8B 权重和 128 个 token 的输入,Odin 在单个 NVIDIA H100 80 GB GPU 上评估了全部 32 个 Transformer 层。服务端端到端 FHE 评估耗时 366.4 秒,峰值设备内存为 58.9 GiB。在相同模型、输入、CKKS 参数和硬件条件下,THOR 耗时 1651.9 秒,实现了 4.51 倍加速。

英文摘要

Cloud LLM services typically require users to send prompts to a model provider, creating a privacy risk. Fully homomorphic encryption (FHE) lets a server perform inference without decrypting the input, but representing data as ciphertexts adds storage and computational overhead. In CKKS-based LLM inference, the packing scheme maps logical tensors to ciphertexts and slots. It therefore determines the ciphertext count and the homomorphic cost of linear layers, and it constrains how data pass between linear layers, attention, and nonlinear computation. As models and sequences grow, inefficient layouts accumulate encoding, compute, and layout-conversion overhead. We present Odin, an FHE inference system that co-designs ciphertext packing and model execution for Llama. Starting from a THOR-style baseline whose bottleneck is weight encoding, Odin uses a feature-major cross-layer layout to unify residual connections and layer interfaces, and builds transient intra-operator layouts for linear projections and attention. This reduces redundant plaintext encoding of weights in wide projections. Within attention, QK^T produces scores that Softmax can consume directly, and PV consumes the resulting probabilities, avoiding intermediate repacking. For nonlinear ops, we use minimax polynomial approximation with input-range control and joint error allocation guided by model quality, reducing polynomial degree and multiplicative depth. To our knowledge, Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3. With Llama-3-8B weights and a 128-token input, Odin evaluates all 32 Transformer layers on a single NVIDIA H100 80 GB GPU. Server-side end-to-end FHE evaluation takes 366.4 s and 58.9 GiB peak device memory. Under the same model, input, CKKS parameters, and hardware, THOR takes 1651.9 s, a 4.51x speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑