用于量子模拟的异构分布式架构
A Heterogeneous Distributed Architecture for Quantum Simulation
浏览论文内容
中文总结 AI 辅助
本文提出一种异构分布式量子模拟架构,适配费米子量子模拟,经450逻辑量子比特的费米-哈伯德与稀疏SYK模型模拟验证,其性能优于同构架构。
中文摘要 AI 辅助
架构专业化与分布式可助力容错量子计算机扩展,但也会因通信、路由及资源复制引入大量开销。本文提出一种异构分布式架构,其中一个魔核(magic core)连接至可扩展存储系统,该系统由专用冷存储节点的一维通道构成,支持对泡利串(Pauli string)奇偶性的并行随机访问。这种组织形式特别适用于费米子量子模拟,可实现这类系统产生的高度非局部泡利串的并行执行。我们在容错模拟中评估该架构,模拟对象为费米-哈伯德(Fermi-Hubbard)模型与稀疏萨赫夫-叶-基塔耶夫(Sachdev-Ye-Kitaev, SYK)模型的动力学,系统规模最大达450个逻辑量子比特。这些工作负载呈现互补的通信结构:费米-哈伯德模型产生由晶格几何决定的从局部到非局部的相互作用谱,而稀疏SYK模型产生高度非局部且重叠的泡利算子。对于450个逻辑量子比特的费米-哈伯德工作负载的一个Trotter步,拥有6条通道和30个T态工厂(T-state factory)的系统,其挂钟时间约为具有4倍数量T态工厂、且具备更强连接性和魔态注入位点的同构分布式架构的1.4倍;在T态工厂数量匹配时,我们的架构速度约快2倍。
英文摘要
Architectural specialization and distribution can help scale fault-tolerant quantum computers, but may also introduce substantial overheads from communication, routing, and resource duplication. We introduce a heterogeneous distributed architecture in which a magic core is connected to an extensible storage system composed of one-dimensional lanes of specialized cold-storage nodes. The storage system supports parallel random access to Pauli string parities. This organization is particularly well suited to fermionic quantum simulation, enabling parallel execution of the highly non-local Pauli strings arising from these systems. We evaluate the architecture on fault-tolerant simulations of the dynamics of the Fermi-Hubbard and sparse Sachdev-Ye-Kitaev (SYK) models on systems of up to 450 logical qubits. These workloads exhibit complementary communication structures: Fermi-Hubbard produces a spectrum of interactions from local to non-local shaped by lattice geometry, whereas sparse SYK produces highly non-local and overlapping Pauli operators. For a Trotter step of a 450-logical-qubit Fermi-Hubbard workload, a six-lane system with 30 T-state factories is within approximately $1.4\times$ the wall-clock time of a homogeneous distributed architecture with 4 times as many T-state factories and substantially greater connectivity and sites for injecting magic. For matched T-factory counts, our architecture is $\sim 2\times$ faster.