arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QSimAdv:一种用于高性能量子电路模拟的晚绑定、厂商无关架构

QSimAdv: A Late-Bound, Vendor-Agnostic Architecture for High-Performance Quantum-Circuit Simulation

Shusen Liu, Pascal Jahan Elahi, Wenyun Sun, Shenjin Lv, Xiaohan Shan, Ugo Varetto

arXiv 2608.16940首次发表:更新:

AI 中文总结

QSimAdv是一种晚绑定、厂商无关的量子电路模拟架构,在NVIDIA GH200等系统上实现,强、弱扩展性优异,性能优于Aer相关配置。

AI 中文摘要

高性能量子电路模拟的可移植性无需从内核层面开始。我们提出QSimAdv,其以晚绑定而非通用内核作为厂商无关性的基础。表示、算子 lower(降级)和数据放置仅在其所需输入可用时才绑定。在全状态分配前,电路、噪声和输出检查可将符合条件的通用采样计数请求路由至stabiliser tableau( stabilizer 表);显式请求的表示保持固定。对于全状态执行,后端约束会形成融合;有序融合算子仅在其物理目标已知后才绑定至原生 lower。一流的逻辑到物理布局图记录了本地和秩地址位上的非规范顺序,因此调度器仅在需要时移动非本地目标。GPU、CPU和消息传递接口(MPI)后端共享这些语义,同时保留原生执行路径。我们在NVIDIA GH200和AMD MI250X/EPYC系统上实现此设计,涵盖本地和分布式执行。在匹配的32位复数浮点状态存储下,QSimAdv在N=32时领先两种Aer Hopper配置,在N=24至30的四种共享MI250X规模下领先Aer的HIP后端。强扩展性暴露了平台依赖性:在setonix上,QSimAdv在每个测量秩上均领先GPU和CPU对比,从1到8个秩分别实现3.4倍和2.8倍的加速;GH200路径和CPU路径在8个秩时均未加速。弱扩展性达到256个秩,GPU状态为2 TiB,CPU状态为1 TiB。这些结果共同表明,可移植性可位于内核边界之上,同时执行保持原生并扩展至分布式内存。

英文摘要

Portability in high-performance quantum-circuit simulation need not begin at the kernel. We present QSimAdv, which makes late binding, rather than a common kernel, the basis of vendor independence. Representation, operator lowering, and data placement are bound only when their required inputs become available. Before full-state allocation, circuit, noise, and output inspection can route eligible generic sampled-count requests to a stabiliser tableau; explicitly requested representations remain fixed. For full-state execution, backend constraints shape fusion; an ordered fused operator binds to a native lowering only after its physical targets are known. A first-class logical-to-physical layout map records non-canonical order across local and rank-address bits, so the dispatcher moves nonlocal targets only on demand. GPU, CPU, and Message Passing Interface (MPI) backends share these semantics while retaining native execution paths. We realize this design on NVIDIA GH200 and AMD MI250X/EPYC systems across local and distributed execution. With matched complex 32-bit floating-point state storage, QSimAdv leads both Aer Hopper configurations at $N=32$ and Aer's HIP backend at four shared MI250X sizes from $N=24$ to 30. Strong scaling exposes platform dependence: on setonix, QSimAdv leads both GPU and CPU comparisons at every measured rank, achieving $3.4\times$ and $2.8\times$ speedups, respectively, from one to eight ranks; neither the GH200 path nor the CPU path speeds up at eight ranks. Weak scaling reaches 256 ranks with 2 TiB GPU and 1 TiB CPU states. Together, these results support that portability can reside above the kernel boundary while execution remains native and extends across distributed memory.

Comments24 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑