arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

近内存处理下量子低密度奇偶校验码的高通量归一化最小和置信传播解码

High-Throughput Normalized Min-Sum Belief Propagation Decoding for Quantum LDPC Codes with Near-Memory Processing

Jeonggeun Seo, Youngsun Han, Leanghok Hour, Dongmin Kim

arXiv 2608.27901首次发表:更新:

AI 中文总结

该研究将qLDPC码的归一化最小和BP解码映射到DPU基PIM架构,实现8.8倍吞吐量提升,且延迟满足量子纠错要求,为qLDPC解码提供高效近内存处理方案。

AI 中文摘要

实时量子纠错需要经典解码器以低且可预测的延迟处理不断增长的校正子工作负载。对于量子低密度奇偶校验(qLDPC)码,迭代置信传播(BP)会在稀疏Tanner图上反复更新消息,带来大量内存访问和数据移动需求。我们将[[144,12,12]]双变量自行车qLDPC码的归一化最小和BP解码映射到基于DPU的内存处理(PIM)架构。每个DPU内,11个任务单元协同解码一个校正子,同时多个DPU并行处理独立的校正子实例。使用uPIMulator和具有理想校正子测量的数据量子比特Pauli误差模型,我们对比了吞吐量、单校正子处理时间、逻辑错误率(LER)和单校正子尾部延迟与16个逻辑CPU基线的表现。在分量级物理错误概率p=0.001且1次BP迭代的情况下,2560个DPU的预计总内核吞吐量达到1.071×10^7次解码/秒,而CPU的吞吐量为1.22×10^6次解码/秒,提升了8.8倍。从2次迭代开始,在所有评估的p值下,测得的LER均低于物理错误概率。在1至5次迭代中,采样的最大序列化X+Z DPU计算延迟仍低于囚禁离子量子纠错的1毫秒解码器侧参考值,5次迭代时约为0.873毫秒。这些结果表明,在评估条件下,近内存处理可为qLDPC的BP解码提供高总吞吐量和亚毫秒级计算延迟。

英文摘要

Real-time quantum error correction requires classical decoders to process growing syndrome workloads with low and predictable latency. For quantum low-density parity-check (qLDPC) codes, iterative belief propagation (BP) repeatedly updates messages over sparse Tanner graphs, creating substantial memory-access and data-movement demands. We map normalized Min-Sum BP decoding of the [[144,12,12]] Bivariate Bicycle qLDPC code onto a DPU-based Processing-in-Memory (PIM) architecture. Within each DPU, 11 tasklets cooperatively decode one syndrome, while multiple DPUs process independent syndrome instances in parallel. Using uPIMulator and a data-qubit Pauli error model with ideal syndrome measurements, we compare throughput, per-syndrome processing time, logical error rate (LER), and single-syndrome tail latency against a 16-logical-CPU baseline. At a component-wise physical error probability of p=0.001 and one BP iteration, the projected aggregate kernel throughput of 2,560 DPUs reaches 1.071 x 10^7 decodes/s, compared with 1.22 x 10^6 decodes/s for the CPU, an 8.8x improvement. From two iterations onward, the measured LER remains below the physical error probability for every evaluated value of p. For one to five iterations, the maximum sampled serialized X+Z DPU compute latency remains below the 1 ms decoder-side reference for trapped-ion QEC, reaching approximately 0.873 ms at five iterations. These results show that near-memory processing can provide high aggregate throughput and sub-millisecond compute latency for qLDPC BP decoding under the evaluated conditions.

Comments22 pages, 8 figures, 1 table. Submitted to Physica Scripta

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑