发表机构
Department of Electrical Engineering, Lahore University of Management Sciences(拉合尔管理科学大学电气工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在商用 ARM Cortex-M33 上实现 ISA 加速且加固的 RBLWE 加密,通过 SIMD 打包与硬件 TRNG 获得最高 4.06 倍加速,并叠加四层对策,验证了加速与安全的实用权衡。
AI 中文摘要
在资源受限的物联网(IoT)设备上实现高效的后量子密码学,需要利用目标处理器架构并抵御实际实现攻击的实现方案。本文在商用 ARM Cortex-M33 微控制器上提出了一种 ISA 加速且实现加固的环二元学习与错误(RBLWE)加密实现。通过将四个 8 位多项式系数打包到 32 位寄存器的字节通道中,并使用 SIMD 风格指令进行处理,配合打包消息编解码器,加速了加密和解密过程,同时从片上安全元件提取的缓冲硬件 TRNG 熵源驱动密钥生成增益。这些共同使得在相同核心冷启动下,密钥生成、加密和解密相对于标量基线分别获得 $4.06\ imes$、$3.35\ imes$ 和 $3.01\ imes$ 的加速,并且加密/解密周期数比参考 Cortex-M0 实现低 $3.52\ imes$/$3.18\ imes$。在此加速核心之上,我们添加了四层分阶段对策:恒定时间执行、针对清零、随机破坏和指令跳过故障的故障加固、Fujisaki-Okamoto(FO)风格的 CCA2 变换,以及一阶共享(掩码)CPA 解密,并单独报告每层的成本。二进制级检查确认这些对策在编译后仍然有效,并识别出一个编译器引入的掩码缺陷,通过手写汇编替换解决。Dudect 风格时序测试、调试器辅助故障注入活动以及组件级 TVLA 为分阶段保护提供了实现级证据。结果表明,在现成微控制器上 RBLWE 具有实用的加速-安全权衡,并为轻量级后量子实现提供了可复用的架构感知技术。
英文摘要
Efficient post-quantum cryptography on resource-constrained Internet-of-Things (IoT) devices requires implementations that exploit the target processor architecture while resisting practical implementation attacks. This paper presents an ISA-accelerated and implementation-hardened realization of Ring Binary Learning with Errors (RBLWE) encryption on a commodity ARM Cortex-M33 microcontroller. Packing four 8-bit polynomial coefficients into the byte lanes of a 32-bit register and processing them with SIMD-style instructions, together with a packed message codec, accelerates encryption and decryption, while a buffered hardware-TRNG entropy source drawn from the on-die Secure Element drives the key-generation gain. Together these give same-core cold-start speedups of $4.06\times$, $3.35\times$, and $3.01\times$ for key generation, encryption, and decryption over a scalar baseline, and $3.52\times$/$3.18\times$ lower encryption/decryption cycle counts than a reference Cortex-M0 implementation. On top of this accelerated core, we add four staged countermeasures: constant-time execution, fault hardening against zeroing, random-corruption, and instruction-skip faults, a Fujisaki-Okamoto (FO)-style CCA2 transform, and first-order shared (masked) CPA decryption, reporting each layer's cost individually. Binary-level inspection confirms these countermeasures survive compilation and identifies a compiler-induced masking flaw resolved with a hand-written assembly replacement. Dudect-style timing tests, debugger-assisted fault-injection campaigns, and component-level TVLA then provide implementation-level evidence for the staged protections. The results demonstrate a practical acceleration-security tradeoff for RBLWE on off-the-shelf microcontrollers and reusable architecture-aware techniques for lightweight post-quantum implementations.