发表机构
CRS4 - Center for Advanced Studies, Research and Development in Sardinia(撒丁岛高级研究、开发与研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将迭代式混合量子-经典优化循环建模为科学工作流,提出终止谓词、子问题级恢复、QPU到经典副本故障切换及统一来源模式四项增强,显著降低编排开销并提升故障韧性。
AI 中文摘要
当今的量子处理单元(QPU)体积过小且噪声过大,无法直接解决大规模组合优化问题,因此实际的混合求解器会将问题分解为多个部分,并在可用的后端(经典启发式算法、模拟器、仿真器或QPU)上迭代执行分解-求解-聚合循环。在实践中,该循环是一个驱动程序脚本,位于量子-HPC中间件之上,负责任务生成、来源追踪、恢复和可移植性。相反,我们将该循环视为一个科学工作流,并探讨工作流层能为通用工作流管理系统和QPU共享中间件增添什么价值。来自实际应用中的两种分解模式——迭代共识(ADMM)和层次划分——对编排层的压力截然不同:在超过120次受管运行中,对于迭代模式,编排耗时占端到端时间的78.4%,其中几乎全部时间花在每轮屏障上;而对于层次模式,编排耗时仅占6.2%。我们的工作流模型增添了通用引擎所不具备的四项功能:驻留在任务图中的终止谓词,用于读取上一轮的残差;子问题级恢复,支持热启动和法定人数延迟聚合;在轮次内从QPU故障切换到经典副本;以及一个来源模式,用相同字段描述CPU、模拟器和QPU,包括所包含的射击预算。我们报告了引擎上每一层的成本,并展示了推测性重新执行如何使运行在注入故障(这些故障会使未受管驱动程序停滞)下仍能完成。我们还利用相同的来源数据,给出了跨模拟器、仿真器和IQM QPU的每设备延迟尾部。
英文摘要
Today's Quantum Processing Units (QPUs) are too small and too noisy to solve large combinatorial optimization problems directly, so practical hybrid solvers split a problem into pieces and iterate a decompose-solve-aggregate loop over whatever backends are available: classical heuristics, simulators, emulators, or a QPU. In practice, the loop is a driver script. It sits on top of the quantum-HPC middleware, handling task generation, provenance, recovery, and portability. Instead, we treat the loop as a scientific workflow and ask what a workflow layer adds to a generic workflow management system and QPU-sharing middleware. Two decomposition patterns from real applications, iterative consensus (ADMM) and hierarchical partitioning, turn out to stress the orchestration layer very differently: over 120 managed runs, orchestration took 78.4% of end-to-end time for the iterative pattern, almost all of it in a per-round barrier, but only 6.2% for the hierarchical one. Our workflow model adds four things a generic engine does not have: a termination predicate residing in the task graph that reads the previous round's residuals, subproblem-level recovery with warm-start and quorum-deferred aggregation, failover from a QPU to a classical replica within a round, and a provenance schema that describes CPUs, simulators and QPUs with the same fields, including the shot budget included. We report the cost of each layer on our engine, show that speculative re-execution enable runs to complete under injected failures that stall an unmanaged driver. We also use the same provenance to give per-device latency tails across simulators, emulators and IQM QPUs.