arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09767cs.CR

使用截断Montgomery技术的更快PMNS多精度乘法

Faster PMNS Multi-precision Multiplications Using Truncated Montgomery Technique

  • Université de Toulon(土伦大学)

机构由 AI 辅助整理,请以论文原文为准。

Laurent-Stéphane Didier, Alexy Dutois-Ruiz, Jean-Marc Robert

AI总结:

本文提出改进的PMNS多精度乘法实现,通过优化参数选择降低内存并提升速度,并应用截断Montgomery技术于内部约简,在AVX512下实现最高15%的模乘加速。

AI中文摘要:

多项式模数系统(PMNS)旨在使用满足某些参数(次数、系数绝对值)界约束的多项式来表示大特征域或环中的元素。当满足某些条件时,通过使用合适的参数和实现特性,这些PMNS可以在密码计算中实现加速。近期工作(Meloni等人\ref{MeloniPV25})通过改进系统的参数生成和使用多精度多项式系数(即使用多个机器字存储的系数)提出了更好的加速。在本工作中,我们首先提出一种PMNS教科书式软件实现,在某些情境下优于\ref{MeloniPV25}的最先进的Toeplitz对应方案,其优势在于通过自由选择更小的$n$参数来降低内存成本,同时最小化乘法复杂度。该实现针对4096位模数大小,元素大小减小11%,速度提升20%。随后,我们对PMNS提出一项新改进,将类似Didier等人\ref{DidierEGR24}提出的截断Montgomery约简技术的方法应用于\textsl{内部约简}。在使用\texttt{AVX512}指令集扩展且模数大小高达8192位的软件实现情境中,与传统方法相比,该新方法在模乘计算中可实现高达15%的加速。

英文摘要:

The Polynomial Modular Number Systems (PMNS) aim to represent elements of fields or rings of large characteristics using polynomials satisfying bounds on some parameters (degree, absolute values of the coefficients). Those PMNS, while some conditions are fulfilled, using convenient parameters and implementation features, allow some speed-ups in cryptographic computations. Recent works (Meloni \emph{et al.} \cite{MeloniPV25}) propose better speed-ups by improving the parameter generation of the system, and by using multi-precision polynomial coefficients, i.e. coefficients stored using several machine words. In this work, we first present PMNS schoolbook software implementation better in some context than the Toeplitz state-of-the-art counterpart of \cite{MeloniPV25}, taking advantage of smaller memory cost by being free to choose a smaller $n$ parameter, minimizing at the same time the complexity of the multiplication. This implementation for 4096 bit modulo size shows element size smaller by 11\% and 20 \% speed-up. We thus present a new improvement on the PMNS, applying to the \textsl{internal reduction} an approach similar to the truncated Montgomery reduction technique presented by Didier \emph{et al.} in \cite{DidierEGR24}. In the context of software implementations using \texttt{AVX512} instruction set extension, and modulo size up to 8192 bits, this new approach allows speed-ups up to 15\% in modular multiplication computation, in comparison with conventional approaches.

补充信息

↑