arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16898cs.ARcs.CRcs.LG

OptiPrime:通过协议-硬件协同设计优化私有推理

OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design

发表机构北京大学 · 开放安全研究院 · 清华大学
查看机构详情
  • Peking University(北京大学)
  • Open Security Research(开放安全研究院)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiangrui Yu, Ye Yu, Si Chen, Chenqi Lin, Wenxuan Zeng, Junfeng Fan, Mingyu Gao, Meng Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对HE-MPC私有推理的网络通信瓶颈,提出协议-硬件协同优化框架OptiPrime,通过新HE卷积协议减少密文传输,并采用权重压缩和专用数据流,实现最高5.7倍(CPU)和4.2倍(加速器)加速。

中文摘要 AI 辅助

基于混合同态加密(HE)和多方计算(MPC)的私有深度神经网络(DNN)推理能够以形式化保证保护用户数据,但代价是由于HE带来的显著延迟开销。定制化的HE加速器已被提出,并在单个HE操作上实现了数量级的加速。然而,当直接将商业HE加速器应用于最先进的HE-MPC框架时,我们观察到端到端性能提升有限。这是因为HE-MPC框架通常需要对每个HE操作的输入和输出密文进行无线传输,导致严重的网络通信瓶颈。为克服这一挑战,我们引入了OptiPrime,一个用于高效私有DNN推理的协议-硬件协同优化框架。OptiPrime提出了一种新颖的用于卷积的HE协议,该协议大幅减少了传输的输出密文数量,并缓解了网络通信瓶颈。同时,由于新协议为更少的输出密文引入了复杂计算,我们观察到由于大量的权重明文和中间密文而带来的新的内存访问挑战。因此,我们进一步提出了一种针对权重明文的轻量级压缩系统,将内存流量减少了10倍,以及一种专门的数据流以最大化中间密文的片上数据重用。大量实验表明,我们的框架在CPU上比Cheetah基线最多快5.7倍,在配备加速器时最多快4.2倍。

英文摘要

Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup for individual HE operations. However, when directly applying a commercial HE accelerator to state-of-the-art HE-MPC frameworks, we observe only limited end-to-end performance gain. This is because HE-MPC frameworks often require wireless transmission of input and output ciphertexts for each HE operation, leading to a severe network communication bottleneck. To overcome this challenge, we introduce OptiPrime, a protocol-hardware co-optimization framework for efficient private DNN inference. OptiPrime features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck. Meanwhile, as the new protocol introduces complex computation for fewer output ciphertext, we observe new memory access challenges due to a high volume of weight plaintexts and intermediate ciphertexts. Hence, we further propose a lightweight compression system for the weight plaintexts, reducing memory traffic by 10 times, as well as a specialized dataflow to maximize on-chip data reuse of intermediate ciphertexts. Extensive experiments show that our framework outperforms the Cheetah baseline by at most 5.7 times on CPUs and 4.2 times with an accelerator.

补充信息

↑