发表机构
Carleton University; Quantegra Technologies Inc.(卡尔顿大学; Quantegra科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
以ML-KEM在Arm Cortex-M7上的案例,展示超越密码内核的系统级优化(如内存层次、外设集成等),在保持算法和线格式不变下,使封装/解封装周期最多减少74.6%/58.8%,提出两阶段方法论。
AI 中文摘要
近期关于嵌入式后量子密码学的研究主要集中于指令级优化,包括算术内核改进、汇编调优、寄存器分配和指令调度。以Arm Cortex-M7上的模块格密钥封装机制(ML-KEM)为案例,我们考察了内存层次利用、紧耦合内存放置、外设集成、时钟配置和确定性公共数据重用所带来的额外收益。评估从最先进的SLOTHY优化实现出发,覆盖了ML-KEM全部三套参数集。在不修改密码算法或标准化线格式的前提下,无辅助公共状态的评估配置将周期数最多减少2.5%。一个选定的公共数据重用配置将封装和解封装周期分别最多减少74.6%和58.8%。这些结果表明,在算术内核优化之后仍有可观的部署收益,并促使我们提出一种同时考察周围执行系统的两阶段方法论。
英文摘要
Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetic-kernel improvements, assembly tuning, register allocation, and instruction scheduling. Using the Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM) on an Arm Cortex-M7 as a case study, we examine the additional gains available from memory-hierarchy utilization, tightly coupled memory placement, peripheral integration, clock configuration, and deterministic public-data reuse. The evaluation starts from a state-of-the-art SLOTHY-optimized implementation and covers all three ML-KEM parameter sets. Without modifying the cryptographic algorithm or standardized wire formats, the evaluated profiles without auxiliary public state reduce cycles by up to 2.5%. A selected public-data-reuse profile reduces encapsulation and decapsulation cycles by up to 74.6% and 58.8%, respectively. These results demonstrate that substantial deployment gains remain after arithmetic-kernel optimization and motivate a two-stage methodology that also examines the surrounding execution system.
Comments22 pages and 4 figures. Extended version to the paper in the Proceedings of FPS-2026: 19th International Symposium on Foundations & Practice of Security