arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2305.11423cs.CRcs.AR

Strix:一种面向具有可编程自举的全同态加密的端到端流式架构,采用两级密文批处理

Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable Bootstrapping

Adiwena Putra, Prasetiyo, Yi Chen, John Kim, Joo-Young Kim

更新

AI总结:

本文提出硬件TFHE加速器Strix,通过两级密文批处理和全流水线FFT微架构解决盲旋转碎片化导致的GPU瓶颈,吞吐量比现有TFHE加速器提升7.4倍。

AI中文摘要:

同态加密(HE)通过将信息隐藏在噪声下以保障安全,从而实现对加密数据的计算。然而,重置密文中噪声级别的自举过程计算成本高昂,且需要庞大的自举密钥。TFHE方案提供了一种更快且可编程的自举算法PBS,这对于机器学习等注重安全的应用至关重要。尽管如此,当前的TFHE方案缺乏对密文打包的支持,导致吞吐量低下。本研究深入分析了TFHE自举,指出了由盲旋转碎片化问题导致的GPU瓶颈,并提出了一种名为Strix的硬件TFHE加速器。Strix引入了两级批处理方法以提升PBS的批处理规模,利用专用微架构实现高效的流式数据处理,并集成了全流水线FFT微架构以提升性能。它在CPU和GPU上的吞吐量均显著高于现有最先进的实现,性能比现有TFHE加速器高出7.4倍。

英文摘要:

Homomorphic encryption (HE) enables computations on encrypted data by concealing information under noise for security. However, the process of bootstrapping, which resets the noise level in the ciphertext, is computationally expensive and requires a large bootstrapping key. The TFHE scheme offers a faster and programmable bootstrapping algorithm called PBS, crucial for security-focused applications like machine learning. Nevertheless, the current TFHE scheme lacks support for ciphertext packing, resulting in low throughput. This work thoroughly analyzes TFHE bootstrapping, identifies the bottleneck in GPUs caused by the blind rotation fragmentation problem, and proposes a hardware TFHE accelerator called Strix. Strix introduces a two-level batching approach to enhance the batch size in PBS, utilizes a specialized microarchitecture for efficient streaming data processing, and incorporates a fully-pipelined FFT microarchitecture to improve performance. It achieves significantly higher throughput than state-of-the-art implementations on both CPUs and GPUs, outperforming existing TFHE accelerators by a factor of 7.4.

↑