arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Python在前,后端派对:跨CPU、GPU和FPGA编译量子工作负载

Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs

Joseph K. L. Lee, Mehrdad Malekmohammadi, Hong-Sheng Zheng, Shuli Shu, Cheick Doumbia, Kalman Szenes, Mehran Zamani Abnili, Thomas Ainsworth, Matthew Seymour, Thomas Germain, Leonhard Neuhaus, Josh Izaac, Lee J. O'Riordan

arXiv 2609.09270首次发表:更新:

发表机构

Xanadu Quantum Technologies Inc.(Xanadu量子技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对量子工作负载跨异构硬件执行的低延迟需求,提出基于PennyLane和Catalyst的Backline编译框架,通过Python前端和MLIR实现CPU、GPU、FPGA的编译执行,实测微秒级往返延迟。

AI 中文摘要

从量子研发迈向生产级、容错的量子工作负载执行,仍然是量子平台构建者面临的最重大挑战之一。虽然Python框架为量子算法设计提供了便捷的入口,但实时量子纠错(QEC)的低延迟要求对性能的需求是传统解释型环境无法满足的。FPGA和ASIC在这些层面发挥着核心作用,但其专门的编程模型使得开发变得僵化且耗时。CPU、GPU和其他加速器引入了不同的挑战:随着基础设施日益异构化,跨不同设备及其相关抽象进行编程变得更加复杂。允许研究人员用高级语言编写工作负载,并将其映射到跨多样分布式目标平台的低延迟执行,将有助于开发实用规模量子系统的关键基础设施。为此,我们引入了Backline,一个在PennyLane和Catalyst中构建的异构编译和运行时框架。Backline使我们能够为高性能和低延迟设备设计和构建量子-经典工作负载,通过MLIR直接从Python接口进行编译。我们演示了多个量子工作负载的编译和执行,这些工作负载在CPU、GPU和FPGA的混合环境中实现低延迟数据传输,适用于本地和分布式远程硬件目标,全部来自与供应商无关的Python前端。以AMD VPK120 FPGA板作为控制器,从其硬件握手引擎发出每一轮,我们测量了通过RoCE v2的稳态中位往返延迟,到AMD Ryzen Threadripper PRO CPU为2.305微秒,到AMD Instinct MI210 GPU为4.5微秒,每个路径执行10^6-1轮,展示了微秒级的同步协同处理。

英文摘要

Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time quantum error correction (QEC) demand performance that traditional interpreted environments cannot provide. FPGAs and ASICs play a central role at these layers, but their specialized programming models make development rigid and time-consuming. CPUs, GPUs, and other accelerators introduce a different challenge: as infrastructure becomes increasingly heterogeneous, programming across different devices and their associated abstractions becomes more complex. Allowing researchers to write workloads in high-level languages that map to low-latency execution across diverse distributed target platforms will enable the development of key infrastructure for utility-scale quantum systems. For this, we introduce $\textit{Backline}$, a heterogeneous compilation and runtime framework built within PennyLane and Catalyst. Backline allows us to design and build quantum-classical workloads for high-performance and low-latency devices, with compilation directly from a Python interface through MLIR. We demonstrate the compilation and execution of several quantum workloads with low-latency data movement across a mix of CPUs, GPUs, and FPGAs, for both local and distributed remote hardware targets, all from a vendor-agnostic Python frontend. With an AMD VPK120 FPGA board as the controller, issuing each round from its hardware-handshake engine, we measured median steady-state round-trip latencies over RoCE v2 of $2.305~μ$s to an AMD Ryzen Threadripper PRO CPU and $4.5~μ$s to an AMD Instinct MI210 GPU across $10^6-1$ rounds per path, demonstrating microsecond-scale synchronous co-processing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑