发表机构
University of Southern California; DEVCOM Army Research Office(南加州大学; 美军研发司令部陆军研究办公室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全同态加密中密钥切换瓶颈,提出自适应选择HKS与KLSS的FPGA加速器,基于性能模型动态切换,在Alveo U280上实现1.84-3.31倍引导加速和1.66-2.52倍图像分类加速。
AI 中文摘要
全同态加密(FHE)能够实现隐私保护的云服务,但会带来巨大的计算开销,因此硬件加速至关重要。在FHE操作中,密钥切换是主要的性能瓶颈。最近的密码学进展提出了一种新颖的密钥切换方法(即KLSS),该方法降低了某些操作复杂度,但比传统的混合密钥切换(HKS)方法需要更高的计算精度。这种权衡导致了不同的计算和内存需求,使得KLSS和HKS的相对延迟高度依赖于硬件并行度、FHE安全参数以及可用的片上内存容量,尤其是在FPGA平台上,内存资源和并行度必须仔细平衡。在这项工作中,我们首先提出了一种内存高效的KLSS数据通路,消除了片外密文传输。然后,我们开发了一个性能模型来分析和比较KLSS和HKS的开销。我们的分析表明,支持这两种方法的自适应解决方案在FHE计算过程中可以实现比静态方法更低的总体延迟。在性能模型的指导下,我们设计了一种基于FPGA的自适应FHE加速器,在计算过程中动态选择HKS和KLSS。我们在Alveo U280上实现了该加速器,并在多个FHE基准上进行了评估。实验结果表明,与最先进的FPGA加速器相比,我们的自适应解决方案在引导(bootstrapping)延迟方面实现了1.84-3.31倍的加速,在安全图像分类方面实现了1.66-2.52倍的加速。
英文摘要
Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential. Among FHE operations, key-switching is a major performance bottleneck. Recent cryptographic advances introduce a novel key-switching method (i.e., KLSS) that reduces certain operational complexity but demands higher computational precision than the traditional Hybrid Key Switching (HKS) method. This trade-off leads to distinct computation and memory requirements, making the relative latency of KLSS and HKS highly dependent on hardware parallelism, FHE security parameters, and available on-chip memory capacity, particularly on FPGA platforms, where memory resources and parallelism must be carefully balanced. In this work, we first propose a memory-efficient KLSS datapath that eliminates off-chip ciphertext transfers. We then develop a performance model to analyze and compare the overheads of both KLSS and HKS. Our analysis reveals that an adaptive solution supporting both methods can achieve lower overall latency than a static method during FHE computation. Guided by the performance model, we design an adaptive FPGA-based FHE accelerator that dynamically selects between HKS and KLSS during computation. We implement the accelerator on an Alveo U280 and evaluate it across multiple FHE benchmarks. Experimental results demonstrate that our adaptive solution achieves a 1.84-3.31$\times$ speedup in bootstrapping latency and a 1.66-2.52$\times$ speedup in secure image classification compared to state-of-the-art FPGA accelerators.