AI 中文总结
本文针对全同态加密效率问题,提出基于Numba的CUDA-Python库LibFHE,重新审视非RNS CKKS-BGV框架并优化,实现高级可编程性与裸机GPU性能,实验证明其性能可媲美高度优化的C++ FHE库,还降低了实现复杂度。
AI 中文摘要
自第四代全同态加密(FHE)框架CKKS提出至今已有十年,却仍无第五代继任者的迹象,近年来众多研究探索GPU加速以提高同态计算效率。本文提出LibFHE,这是一个高性能的GPU加速框架,通过CUDA-Python绑定,为同态工作负载实现高级可编程性和裸机GPU性能。多数现有实现采用RNS-CKKS变体,本文特意重新审视原始(非RNS)CKKS-BGV框架并开发基于GPU的实现及相应优化。实验结果表明,优化后的CUDA-Python实现可达到与高度优化的C++ FHE库相当的性能,同时显著降低实现复杂度并提高可编程性。
英文摘要
It has been a decade since the fourth-generation FHE framework, CKKS, was proposed; yet, there is still no indicator pointing toward a fifth-generation successor; and in recent years, numerous studies have explored GPU acceleration to improve the efficiency of homomorphic computations. In this paper, we propose LibFHE, a high-performance GPU-accelerated framework that features CUDA-Python bindings to achieve both high-level programmability and bare-metal GPU performance for homomorphic workloads. A large majority of state-of-the-art implementations adopt the RNS-CKKS variant. In contrast, this work deliberately revisits the original (non-RNS) CKKS-BGV framework, and develops a GPU-based implementation along with corresponding optimizations. Experimental results demonstrate that optimized CUDA-Python implementations can achieve performance comparable to highly optimized CPU-based C++ FHE libraries, while significantly reducing implementation complexity and improving programmability.