arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于后量子密码学的带错误学习密钥封装机制的便携式加速

Portable Acceleration of Learning With Errors KEMs for Post-Quantum Cryptography

Tiziana Liberati, Nitin Shukla, Simone Rizzo, Elisabetta Boella, Matteo Barbieri, Gabriella Bettonte, Daniele Gregori, Marco Pedicini

arXiv 2607.09541首次发表:更新:

AI 中文总结

研究针对后量子密码学中基于带错误学习的密钥封装机制计算成本高的问题,提出基于OpenMP Target卸载的便携式GPU实现方法,经多方面评估,该方法能显著加速且保持可移植性,降低计算开销,避免供应商锁定。

AI 中文摘要

向量子后密码学(PQC)的转变推动了对能满足实际应用计算需求的实现的需求。在提议的PQC构造中,基于带错误学习(LWE)的密钥封装机制(KEM)因其强大的安全基础而特别有吸引力,但矩阵运算和大规模密码安全随机数生成会带来大量计算成本。GPU加速是降低基于格的加密方案计算开销的有效方法。本文提出了一种基于OpenMP Target卸载的便携式GPU实现的普通LWE基KEM。与大多数依赖CUDA特定优化的现有GPU实现不同,我们的方法使用在NVIDIA和AMD加速器上都能执行的单一源代码库。我们在不同加速器架构上评估了该实现,分析了性能基准测试、运行时剖析、可扩展性分析和能量消耗测量。实验结果表明,OpenMP Target卸载在多核CPU基准上实现了显著加速,同时在异构GPU生态系统中保持了源级可移植性。跨平台分析确定NVIDIA GH200和AMD MI300X是这种内存受限工作负载最有效的平台,剖析表明内存系统组织和CPU-GPU交互比单独的峰值计算能力起更关键的作用。这些发现表明便携式GPU加速可以显著降低PQC的计算开销,同时避免供应商锁定,从而促进抗量子加密基础设施的部署。

英文摘要

The transition to post-quantum cryptography (PQC) is driving demand for implementations that can meet the computational requirements of real-world applications. Among the proposed PQC constructions, Learning With Errors (LWE) based key encapsulation mechanisms (KEMs) are particularly attractive due to their strong security foundations, but they incur substantial computational costs from matrix operations and large-scale cryptographically secure random number generation. These characteristics position GPU acceleration as an effective approach for lowering the computational overhead of lattice based cryptographic schemes. In this work, we present a portable GPU implementation of a plain LWE based KEM using OpenMP Target offloading. Unlike most existing GPU implementations, which rely on CUDA specific optimizations, our approach uses a single source code base that executes on both NVIDIA and AMD accelerators. We evaluate the proposed implementation on different accelerator architectures, analyzing performance benchmarking, runtime profiling, scalability analysis, and energy to solution measurements. Experimental results show that OpenMP Target offloading delivers substantial acceleration over a multicore CPU baseline while preserving source level portability across heterogeneous GPU ecosystems. Cross platform analysis identifies NVIDIA GH200 and AMD MI300X as the most effective platforms for this memory bound workload, while profiling indicates that memory system organization and CPU GPU interaction play a more critical role than peak compute capability alone. These findings demonstrate that portable GPU acceleration can significantly reduce the computational overhead of PQC while avoiding vendor lock in, thereby facilitating the deployment of quantum resistant cryptographic infrastructures.

DOI:10.5281/zenodo.21297470

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑