arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13948cs.CR

基于AVX-512的向量化SQIsign实现

Exposing SIMD Parallelism in SQIsign: An AVX-512 Implementation

Weize Wang, Chutong Wang, Yu Wu, Qifan Xue, Jieyu Zheng, Yunlei Zhao

AI总结:

本文提出首个利用AVX-512 IFMA指令集的SQIsign向量化实现,结合Qlapoti技术在NIST安全级别I下获2.69倍签名、3.18倍验证提速,为后量子同源密码优化提供方法基础。

AI中文摘要:

SQIsign是提交至NIST后量子密码标准制定流程中唯一基于同源性的数字签名方案,其特点是建立在 supersingular 椭圆曲线的自同态环问题的难解性之上。尽管SQIsign具备紧凑的密钥和签名尺寸,但其实际部署受到计算密集型签名流程的阻碍。本文提出了首个利用AVX-512整数融合乘加(IFMA)指令集架构的SQIsign综合向量化实现。通过系统性重新设计计算栈——涵盖素数域与扩域算术、包含批量点加倍和标量乘法的椭圆曲线运算,以及通过立方算术和二维同源性评估实现的配对计算——我们实现了远超参考实现的性能提升。结合Qlapoti技术,在NIST安全级别I下,我们的实现实现了2.69倍的签名速度提升和3.18倍的验证速度提升。与关于AVX-512已过时的误解相反,我们强调英特尔的AVX10指令集架构(修订版10.2,计划于2026年末广泛部署)将在性能核和效率核上标准化AVX-512能力——包括IFMA指令——确保这些优化技术的长期可行性。此外,我们的向量化策略与架构无关,为更广泛的基于同源性的密码构造提供了可应用的方法学基础。本工作表明,SIMD向量化是后量子同源性方案中一个关键但未被充分探索的优化维度,与近期算法进展无关。

英文摘要:

Modern isogeny-based cryptosystems spend much of their running time in finite-field, elliptic-curve, and higher-dimensional isogeny arithmetic. Exploiting SIMD parallelism is nontrivial: routines such as Montgomery ladders contain loop-carried dependencies, while point, pairing, and theta-coordinate formulas expose only irregular fine-grained parallelism. We show that substantial SIMD parallelism can be recovered by reorganizing the arithmetic dependency graphs of higher-level primitives rather than vectorizing field multiplication in isolation. We develop an end-to-end AVX-512IFMA implementation of SQIsign in which data remain in a radix-$2^{51}$ vector representation across most of the curve-side computation. Our redesign includes projective xDBLADD schedules, batched point doubling in several coordinate systems, a vectorized biscalar ladder, fused cubical-arithmetic pairing steps, and batched one- and two-dimensional isogeny evaluation. Relative to the reference C implementation, we achieve end-to-end speedups of $1.76\times$, $1.71\times$, and $3.18\times$ for key generation, signing, and verification at NIST level~I; combined with Qlapoti, key-generation and signing speedups rise to $2.90\times$ and $2.69\times$. We further apply the same backend and methodology to CORAL, a recent isogeny group action for post-quantum non-interactive key exchange based on two-dimensional $2$-isogenies. Across five parameter sets, this yields $1.28$--$1.40\times$ speedups for key generation and $1.92$--$2.46\times$ for shared-key computation. These results provide cross-scheme evidence that algorithm-level SIMD scheduling is a reusable optimization dimension for higher-dimensional isogeny cryptography.

↑