arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向字级全同态加密的内存高效设计

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Ardhi Wiratama Baskara Yudha, Erwin Eko Wahyudi, Rian Adam Rajagede, Qian Lou, Yan Solihin

arXiv 2609.04769首次发表:更新:

发表机构

Advanced Micro Devices, Inc.; University of Central Florida(超威半导体; 中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FHE优化框架BXT,通过四项技术缓解内存瓶颈,在CNN推理中实现最高3.8倍加速且精度损失小于1%。

AI 中文摘要

全同态加密(FHE)支持对加密数据进行计算,在分析全过程中保护隐私。尽管其隐私性极强,但FHE的执行速度远慢于原始计算。尤其在近期计算加速取得成功后,性能瓶颈转向内存,考虑到FHE会将数据大小放大数个数量级,导致算术强度较低。我们提出BXT,这是一种FHE优化框架,通过四项技术缓解内存瓶颈:(1)密文压缩,在执行期间从种子重新生成密文组件;(2)密文序列化,将系数打包为位数组,并在L2到L1的传输过程中解包;(3)延迟种子生成,在聚合操作间推迟依赖伪随机数生成器(PRNG)的离线工作;(4)面向通用FHE的、由故障感知训练引导的密文数字剪枝。在卷积神经网络(CNN)推理任务中,BXT-CSO50配置实现了较100倍GPU基准最高3.8倍的加速,在50%比较精度下精度损失小于1%。

英文摘要

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the memory, especially considering that FHE magnifies the data size by orders of magnitude, resulting in a low arithmetic intensity. We propose BXT, an FHE optimization framework that mitigates the memory bottleneck through four techniques: (1) ciphertext compression, which regenerates ciphertext components from seeds during execution; (2) ciphertext serialization, which packs coefficients as bit arrays and unpacks them during L2-to-L1 transfer; (3) delayed seed generation, which defers PRNG-heavy offline work across aggregated operations; and (4) ciphertext digit pruning guided by fault-aware training tailored for Universal FHE. On CNN inference, the BXT-CSO50 configuration effectively achieves up to 3.8$\times$ speedup over the 100x GPU baseline with less than 1% accuracy loss at 50% comparison precision.

CommentsAccepted at ICCD 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑