发表机构
National Institute of Technology Srinagar(斯利那加国立理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出APGEM,一种上下文感知的错误缓解编排框架,动态选择ZNE、PEC、CDR或REM策略,在NISQ设备上提升量子强化学习的鲁棒性,在CVRP上达到Oracle策略效用的94%。
AI 中文摘要
量子强化学习(QRL)将强化学习与参数化量子电路相结合,是解决组合优化问题的一种有前景的方法。然而,在含噪声的中等规模量子(NISQ)设备上,退相干、门缺陷和测量误差会降低策略质量,并使学习过程变得不那么可靠。现有的错误缓解技术通常作为固定校正应用,不能适应变化的噪声条件或训练状态的演变。本工作提出自适应策略引导的错误缓解(APGEM),作为混合量子-经典训练循环的上下文感知编排层,在QRL训练期间动态选择最合适的缓解策略。APGEM使用策略级指标(包括量子态保真度、策略熵、累积奖励和近似比)评估零噪声外推(ZNE)、概率误差消除(PEC)、Clifford数据回归(CDR)和读出错误缓解(REM),并将所选策略直接集成到强化学习循环中。该框架在容量车辆路径问题(CVRP)上进行了评估,CVRP是城市物流中一个代表性的NP难问题,在一系列NISQ噪声模型和噪声水平下进行。APGEM始终优于传统的静态缓解方法,达到Oracle策略效用的约94%,在噪声增加时保持更高的量子态保真度,并在整个训练过程中产生更稳定的学习行为。消融研究表明,该框架学习了上下文感知的缓解策略,能够适应不同的噪声环境和电路执行条件。这些发现表明,将自适应错误缓解集成到学习过程中显著提高了QRL在NISQ硬件上的鲁棒性和可靠性。
英文摘要
Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
Comments9 pages, 18 figures, 13 tables