AI 中文总结
该研究针对ARM Cortex-M4优化HQC实现,通过改进多项式乘法、固定权重采样及新增缓存策略,在NUCLEO-L4R5ZI板上大幅降低了HQC的密钥生成、封装与解封装开销。
AI 中文摘要
本文提出了ARM Cortex-M4上HQC(汉明准循环)的优化实现。我们优化了三部分内容:(i)多项式乘法;(ii)固定权重采样中的支撑集扩展;(iii)提出了一种可选的缓存策略,可复用固定密钥下重新计算的公共变换与哈希。对于多项式乘法,Frobenius加法FFT(FAFFT)蝶形运算中的固定常数乘法,近一半指令用于通用寄存器与浮点寄存器间的VMOV数据移动,而非算术运算。由于仅最小化异或计数会增加总指令数,我们提出了感知脏态的寄存器分配策略与异或操作重排序,在保持异或计数不变的同时,将VMOV计数最多降低48.1%。我们将这些优化应用于结合现有FAFFT-CRT方法的乘法,针对HQC-1进一步发现了稀疏度达34%的FAFFT模,降低了CRT重构开销。对于固定权重采样,我们用谓词执行与4路展开重写了支撑集扩展,在保持常数时间的同时,将其内循环的单字开销从22周期降至6周期。在NUCLEO-L4R5ZI开发板上,我们的实现相较于两种现有最先进实现中较快的一种,将密钥生成、封装、解封装分别最多降低33.1%、34.6%、29.8%;可选缓存策略还可进一步将封装、解封装分别最多降低32.7%、18.9%。
英文摘要
In this paper, we present an optimized implementation of Hamming Quasi-Cyclic (HQC) on the ARM Cortex-M4. We optimize (i) the polynomial multiplication and (ii) the support expansion in fixed-weight sampling, and (iii) propose an optional caching strategy that reuses the public transforms and hash recomputed under a fixed key. For the polynomial multiplication, the fixed-constant multiplications in the Frobenius additive FFT (FAFFT) butterfly spend nearly half of their instructions on VMOV data movements between general-purpose and floating-point registers rather than arithmetic. Because minimizing the XOR count alone can increase the total instruction count, we propose a dirty-aware register-allocation policy and an XOR-operation reordering that reduce the VMOV count by up to 48.1% while leaving the XOR count unchanged. We apply these to a multiplication that combines prior FAFFT-CRT methods, and for HQC-1 we further find a 34% sparser FAFFT modulus that lowers the CRT reconstruction cost. For fixed-weight sampling, we rewrite the support expansion with predicated execution and 4-way unrolling, lowering the per-word cost of its inner loop from 22 to 6 cycles while remaining constant-time. On the NUCLEO-L4R5ZI board, our implementation reduces key generation, encapsulation, and decapsulation by up to 33.1%, 34.6%, and 29.8% over the faster of the two prior state-of-the-art implementations, and the optional caching yields a further reduction of up to 32.7% and 18.9% for encapsulation and decapsulation.
Comments22 pages, 1 figure, 11 tables