arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31252cs.CRcs.AR

Peregrino:面向资源受限边缘设备的完整Falcon后量子数字签名方案的全硬件加速器

Peregrino: A Full-Hardware Accelerator for the Complete Falcon Post-Quantum Digital Signature Scheme on Resource-Constrained Edge Devices

发表机构马德里理工大学
查看机构详情
  • Universidad Politécnica de Madrid(马德里理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Antonio Carreño, Jaime Señor, Jorge Portilla

首次发表
浏览论文内容

中文总结 AI 辅助

针对Falcon后量子签名在硬件实现的难题,提出Peregrino加速器,用HDL从零设计完整方案,以模拟浮点适配无FPU边缘设备,资源占用大幅低于HLS方案,性能提升显著。

中文摘要 AI 辅助

量子计算机的到来威胁着经典密码学的安全保证,因为量子算法可以破解那些在传统攻击下仍然安全的方案。因此,美国国家标准与技术研究院(NIST)标准化了一组后量子密码算法,其中就包括Falcon,这是一种基于格的数字签名方案,在标准化候选方案中具有最紧凑的签名和公钥尺寸。Falcon对浮点运算的依赖使其难以在硬件中实现,而先前的工作仅提供了针对特定操作(如签名生成或验证)的部分加速器,或者通过高层次综合(HLS)生成的单个完整实现。本工作提出了Peregrino,这是首个从头开始用HDL设计的完整Falcon数字签名方案的硬件加速器,通过模拟浮点数据路径,面向没有原生浮点单元的资源受限边缘设备。在单个Artix 7 XC7A200T FPGA上实现时,Peregrino在Falcon-1024变体上使用了85261个LUT、41382个FF、44个BRAM和142个DSP。与先前唯一的完整实现(基于HLS的FalconTakesOff)相比,它使用的LUT减少了1.9倍,FF减少了3.4倍,BRAM减少了2.7倍,DSP减少了9.9倍,使得整个方案能够适配在单个FPGA上,而HLS设计至少需要两个FPGA。作为片上MicroBlaze软核的外设,该加速器还将密钥对生成、签名生成和签名验证的时钟周期分别减少了92%、96%和85%,相对于模拟浮点参考软件。

英文摘要

The arrival of quantum computers threatens the security guarantees of classical cryptography, since quantum algorithms can break schemes that remain secure against conventional attacks. The National Institute of Standards and Technology (NIST) has therefore standardized a set of post-quantum cryptographic algorithms, among them Falcon, a lattice-based digital signature scheme with the most compact signature and public-key sizes of the standardized candidates. Falcon's reliance on floating-point arithmetic makes it hard to implement in hardware, and prior work offers only partial accelerators for specific operations such as signature generation or verification, or a single full implementation generated through high-level synthesis (HLS). This work presents Peregrino, the first hardware accelerator of the complete Falcon digital signature scheme designed from scratch in HDL, targeting resource-constrained edge devices without a native floating-point unit through an emulated floating-point datapath. Implemented on a single Artix 7 XC7A200T FPGA, Peregrino uses 85261 LUTs, 41382 FFs, 44 BRAMs, and 142 DSPs for the Falcon-1024 variant. Against the only prior full implementation, the HLS-based FalconTakesOff, it uses 1,9$\times$ fewer LUTs, 3,4$\times$ fewer FFs, 2,7$\times$ fewer BRAMs, and 9,9$\times$ fewer DSPs, fitting the entire scheme on one FPGA where the HLS design requires at least two. Operating as a peripheral of an on-chip MicroBlaze soft-core, the accelerator additionally reduces key-pair generation, signature generation, and signature verification clock cycles by 92\%, 96\%, and 85\% over the emulated floating-point reference software.

↑