PHAT:面向TFHE的光子加速器
PHAT: PHotonic Accelerator for TFHE
- Boston University(波士顿大学)
- University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PHAT是一种利用光学寻址相变存储器的电光加速器,通过优化的FFT单元、旋转因子固定数据流和调度机制,在真实TFHE工作负载上相比最先进ASIC实现2.14至5.10倍加速。
AI中文摘要:
全同态加密(FHE)能够在加密数据上进行安全计算,使其成为云环境中隐私保护应用的一种有前景的解决方案。在各种FHE方案中,基于环面(Torus)的FHE(TFHE)因其支持任意操作而脱颖而出。然而,其高昂的计算和通信开销,特别是在引导(bootstrapping)过程中所需的快速傅里叶变换(FFT)操作,限制了其在现实世界应用中的实用性。传统的电子加速器由于技术缩放的限制和内存墙问题,难以实现足够的吞吐量。为了应对这些挑战,我们提出了PHAT,一种利用光学寻址相变存储器(OPCM)的TFHE光子加速器。基于OPCM的存内计算系统能够提供高计算和通信吞吐量,使其非常适合加速TFHE中的FFT操作。然而,直接将FFT映射到OPCM面临着高精度模拟计算以及OPCM单元编程的高延迟和高能耗等挑战。为了克服这些困难,我们引入了一种针对TFHE优化的新型电光加速器架构,其特点包括基于OPCM的FFT单元、为OPCM量身定制的旋转因子固定(twiddle-stationary)数据流,以及一种最大化FFT单元利用率的调度机制。在四个真实世界的TFHE工作负载上,与最先进的ASIC加速器相比,PHAT实现了2.14倍至5.10倍的加速比。我们的方法显著提升了TFHE应用的性能,为云计算中实用且高效的同态加密铺平了道路。
英文摘要:
Fully Homomorphic Encryption (FHE) enables secure computation on encrypted data, making it a promising solution for privacy-preserving applications in the cloud. Among various FHE schemes, FHE over the Torus (TFHE) stands out due to its support for arbitrary operations. However, its high computation and communication overhead, particularly in the Fast Fourier Transform (FFT) operations required during bootstrapping, limits its practicality for real-world applications. Conventional electronic accelerators struggle to achieve sufficient throughput due to the limitations of technology scaling and the memory-wall problem. To address these challenges, we propose PHAT, a PHotonic Accelerator for TFHE leveraging Optically-addressed Phase-Change Memory (OPCM). OPCM-based processing-in-memory systems offer high computation and communication throughput, making them well-suited for accelerating FFT operations in TFHE. However, directly mapping FFT to OPCM presents challenges such as high-precision analog computation and the high latency and energy cost of programming OPCM cells. To overcome these challenges, we introduce a novel electro-photonic accelerator architecture optimized for TFHE, featuring OPCM-based FFT units, a twiddle-stationary dataflow tailored for OPCM, and a scheduling mechanism to maximize the utilization of the FFT units. PHAT delivers $2.14\times$--$5.10\times$ speedup across four real-world TFHE workloads against the state-of-the-art ASIC accelerator. Our approach significantly enhances the performance of TFHE applications, paving the way for practical and efficient homomorphic encryption in cloud computing.