arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11107cs.CR

面向x86物联网网关的后量子TLS 1.3的HQC加速

Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways

Jihoon Jang, Hyunju Park, Jebin Kim, Seokhie Hong, Suhri Kim

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对x86物联网网关优化HQC算法并集成到TLS 1.3,通过AVX2、AVX-512及GFNI等技术加速,显著提升了HQC的密钥生成、封装、解封装速度,降低了TLS握手延迟,明确了受限链路上HQC握手成本的决定因素。

中文摘要 AI 辅助

物联网网关上的后量子TLS必须在不显著增加握手延迟、CPU开销或网络流量的前提下保护大量设备连接。HQC提供了超越ML-KEM的基于编码的多样性,但其计算开销和密文大小可能会增加这些成本。我们针对带有AVX2、AVX-512和伽罗瓦域新指令(GFNI)的x86处理器优化HQC,并将优化后的实现集成到TLS 1.3中。我们将无分支的Toom-Cook/Karatsuba乘法扩展到AVX-512,用GFNI加速里德-所罗门(Reed-Solomon)解码,改进里德-穆勒(Reed-Muller)解码,加速SHA3-512和固定权重采样,并将弗罗贝尼乌斯加法快速傅里叶变换(FAFFT)乘法移植到AVX-512+GFNI以适配HQC-5。在英特尔酷睿i5-1135G7处理器上,我们的AVX2实现相比现有最快的AVX2结果,将解封装速度提升了11.1%-21.5%;我们的AVX-512实现相比Cabral等人的工作,在三种HQC参数集下,密钥生成速度提升24.9%-29.5%、封装速度提升10.4%-11.1%、解封装速度提升20.6%-26.1%。我们评估了这些实现对TLS 1.3握手延迟、服务器与客户端CPU开销的影响,以及往返时间(RTT)和带宽的影响。在本地回环TLS测量中,我们的AVX-512 HQC-5实现相比Cabral等人的工作,将握手延迟从3.22ms降至2.92ms;在带宽1Mbit/s、额外RTT 50ms的受限路径上,HQC-5握手耗时243ms,而ML-KEM-1024仅为87ms,这表明受限链路上剩余的握手成本由HQC的公钥和密文大小而非实现速度决定。

英文摘要

Post-quantum TLS at an IoT gateway must protect many device connections without large increases in handshake delay, CPU cost, or network traffic. HQC provides code-based diversity beyond ML-KEM, but its computation and ciphertext sizes can increase these costs. We optimize HQC for x86 processors with AVX2, AVX-512, and the Galois Field New Instructions (GFNI), and integrate the resulting implementations into TLS 1.3. We extend branch-free Toom-Cook/Karatsuba multiplication to AVX-512, accelerate Reed-Solomon decoding with GFNI, improve Reed-Muller decoding, accelerate SHA3-512 and fixed-weight sampling, and port Frobenius additive FFT (FAFFT) multiplication to AVX-512 + GFNI for HQC-5. On an Intel Core i5-1135G7 processor, our AVX2 implementation reduces decapsulation by 11.1-21.5% over the fastest prior AVX2 results. Our AVX-512 implementation reduces key generation by 24.9-29.5%, encapsulation by 10.4-11.1%, and decapsulation by 20.6-26.1% relative to Cabral et al. across the three HQC parameter sets. We evaluate the effects of these implementations on TLS 1.3 handshake latency and server and client CPU costs, as well as the effects of round-trip time (RTT) and bandwidth. In local loopback TLS measurements, our AVX-512 HQC-5 implementation reduces handshake latency from 3.22 ms to 2.92 ms compared with Cabral et al. On a constrained path with 1 Mbit/s bandwidth and an added RTT of 50 ms, the HQC-5 handshake takes 243 ms, compared with 87 ms for ML-KEM-1024. This indicates that HQC public-key and ciphertext sizes, rather than implementation speed, determine the remaining handshake cost on constrained links.

发表机构

  • School of Cyber Security, Korea University(韩国大学网络安全学院)
  • School of Mathematics, Statistics and Data Science, Sungshin Women’s University(圣心女子大学数学、统计与数据科学学院)
  • SmartM2M

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑