arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ML-KEM的NTT算术在CGLA上的实现与评估

Implementation and Evaluation of NTT Arithmetic for ML-KEM on a CGLA

Takuto Ando, Yasuhiko Nakashima

arXiv 2609.28904首次发表:更新:

发表机构

Nara Institute of Science and Technology(奈良先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文在CPU-接地线性阵列上实现了ML-KEM的NTT算术,通过拆分旋转因子和融合阶段,在FPGA和ASIC上实现了低延迟与低能耗的精确计算。

AI 中文摘要

FIPS 203标准将ML-KEM用于后量子密钥建立。其多项式乘法依赖于在模数q=3329下进行精确模运算的NTT蝶形运算。专用的NTT加速器通过固定的模运算和阶段调度来最小化延迟。CPU-接地线性阵列(CGLA)在多个工作负载中重用一条可编程的线性数据通路。将该变换映射到该数据通路需要在ARM到CGLA接口上进行精确的FP32重建和显式的阶段转换。我们通过将每个旋转因子在模约简前拆分为8位和4位部分,在ML-KEM模数上实现了一个八级循环基-2驱动程序。该驱动程序不同于标准化的七层不完全负循环NTT,并未实现完整的ML-KEM多项式乘法路径。该算术序列使每个整数保持在2^24以下。一次41-PE调用融合了前两个基-2阶段,六次47-PE调用执行其余阶段,同时将输出分散到下一阶段的记录中。在四个队列中,105次FPGA运行匹配了所有817152个输出系数。在批大小为64时,实测FPGA端到端延迟为每个NTT 27.5微秒,ASIC预测为6.74微秒。通过PE门控,预测的ASIC系统能量在批大小为8时为每个NTT 10.1微焦耳,在批大小为64时为58.9微焦耳。

英文摘要

FIPS 203 standardizes ML-KEM for post-quantum key establishment. Its polynomial multiplication relies on NTT butterflies with exact modular arithmetic over q = 3329. Dedicated NTT accelerators minimize latency with fixed modular arithmetic and stage schedules. A CPU-Grounded Linear Array (CGLA) reuses one programmable linear datapath across several workloads. Mapping the transform to this datapath requires exact FP32 reconstruction and explicit stage transitions across the ARM-to-CGLA interface. We implement an eight-stage cyclic radix-2 driver over the ML-KEM modulus by splitting each twiddle into 8-bit and 4-bit parts before modular reduction. This driver differs from the standardized seven-layer incomplete negacyclic NTT and does not implement the full ML-KEM polynomial multiplication path. The arithmetic sequence keeps every integer below 2^24. One 41-PE call fuses the first two radix-2 stages, and six 47-PE calls execute the remaining stages while scattering outputs into next-stage records. Across four cohorts, 105 FPGA runs match all 817152 output coefficients. At batch size 64, measured FPGA end-to-end latency is 27.5 us per NTT, and the ASIC projection is 6.74 us. With PE gating, projected ASIC system energy is 10.1 uJ per NTT at batch size 8 and 58.9 uJ at batch size 64.

CommentsAccepted as a Regular Paper for the WICS Workshop at CANDAR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑