arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qupertino:纯MLX数组内核与手工调优Metal着色器在Apple Silicon上的量子电路模拟对比

Qupertino: Pure MLX Array Kernels versus Hand-Tuned Metal Shaders for Quantum Circuit Simulation on Apple Silicon

Shlomo Kashani

arXiv 2609.19147首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Qupertino是Apple Silicon上的开源量子电路模拟器,对比纯MLX数组内核与手工Metal着色器,着色器层级在25量子位下实现最高95倍加速,并支持多种工作负载。

AI 中文摘要

我们提出了Qupertino,一个面向Apple Silicon的开源量子电路模拟器,并利用它来探究纯MLX数组操作编写的模拟器能达到何种程度,以及手工调优的Metal着色器还能带来什么改进。该框架提供两个可测量的层级。纯MLX层级将结构化门分派到专门的MLX内核:对角门作为广播相位乘法运行,受控门作为掩码半态更新,SWAP门作为轴置换;一个配对的密集路径消融实验表明,仅此分派就带来了25-33倍的加速。可选的着色器层级为基准测试中的每个结构化层族添加了手写Metal内核,包括相位查找表对角门、GF(2)仿射置换收集、融合张量积单量子位层、基数4 QFT和Walsh-Hadamard蝶形变换,以及基共轭XX/YY Trotter层;运行时融合检测器将工作路由到这些内核,同时通过奇偶校验测试确认精确保持电路语义。在M1 Max上进行的四路交错测试中(每个单元两次预热、十次测量重复),着色器层级在全部18个对比单元中,相对于同机Qiskit Aer CPU和PennyLane的平均运行时间均为最快。在25个量子位下,门流QFT运行时间为0.0591 +/- 0.0029秒(配对加速比分别为Aer的47.2倍和PennyLane的95.3倍),TFIM Trotter演化运行时间为0.495 +/- 0.035秒(加速比分别为36.2倍和67.2倍)。在包含29个工作负载的测试套件中,着色器层级相对于纯MLX的配对加速比达到25倍,其中25个(共29个)工作负载的加速超过1倍。该框架还支持变分拟设工作负载、QAOA、QCBM、Trotter-Suzuki哈密顿量模拟以及OpenQASM 2.0导入(酉子集);一个初步的MPS后端在有限纠缠工作负载上可达150个量子位。正确性依赖于253个Python测试、独立的complex128检查以及针对精确对角化的Trotter误差曲线。态矢量内存仍随量子位数呈指数增长。

英文摘要

We present Qupertino, an open-source quantum circuit simulator for Apple Silicon, and use it to ask how far a simulator written purely in MLX array operations can go and what remains for hand-tuned Metal shaders. The framework ships two measured tiers. The pure tier dispatches structured gates to specialized MLX kernels: diagonal gates run as broadcast phase multiplies, controlled gates as masked half-state updates, and SWAP as an axis permutation; a paired dense-path ablation attributes a 25-33x speedup to this dispatch alone. The opt-in shader tier adds hand-written Metal kernels for every structured layer family in our benchmarks, including phase-LUT diagonals, GF(2) affine permutation gathers, fused tensor-product single-qubit layers, radix-4 QFT and Walsh-Hadamard butterflies, and basis-conjugated XX/YY Trotter layers; runtime fusion detectors route work to them while preserving circuit semantics exactly, confirmed by parity tests. In a four-way interleaved campaign on M1 Max (two warmups, ten measured repeats per cell), the shader tier is fastest by mean runtime in all 18 comparison cells against same-machine Qiskit Aer CPU and PennyLane lightning.qubit. At 25 qubits, gate-stream QFT runs in 0.0591 +/- 0.0029 s (paired 47.2x over Aer, 95.3x over PennyLane) and TFIM Trotter evolution in 0.495 +/- 0.035 s (36.2x and 67.2x). Across a 29-workload suite, the shader tier's paired speedup over pure MLX reaches 25x, with 25 of 29 workloads accelerating above parity. The framework also supports variational ansatz workloads, QAOA, QCBM, Trotter-Suzuki Hamiltonian simulation, and OpenQASM 2.0 import (unitary subset); a preliminary MPS backend reaches 150 qubits on limited-entanglement workloads. Correctness rests on 253 Python tests, independent complex128 checks, and a Trotter-error curve against exact diagonalization. State-vector memory remains exponential in qubit count.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑