arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18219cs.LGcs.AIcs.ET

APGEM:面向真实世界CVRP案例的自适应策略引导的量子强化学习误差缓解

APGEM: Adaptive Policy-Guided Error Mitigation for Quantum Reinforcement Learning on a Real-World CVRP Case Study

Shabir Ahmad Sofi, Bisma Majid, Mir Mohammad Yousuf

首次发表
浏览论文内容

中文总结 AI 辅助

针对量子强化学习在真实CVRP上受噪声影响的问题,提出自适应策略引导的误差缓解控制器APGEM,在线选择ZNE、PEC、CDR和REM,将高噪声下近似比从0.84-0.87提升至0.92-0.94。

中文摘要 AI 辅助

量子强化学习(QRL)将策略表示为变分量子电路(VQCs),这使得它对于诸如带容量约束的车辆路径问题(CVRP)等组合优化问题具有吸引力。然而,在含噪声的中等规模量子(NISQ)硬件上,退相干会降低保真度并破坏学习的稳定性,而传统的误差缓解是静态应用的,不考虑学习上下文。我们引入了自适应策略引导的误差缓解(APGEM),这是一个控制器,它在线选择零噪声外推(ZNE)、概率误差消除(PEC)、克利福德数据回归(CDR)和读出误差缓解(REM),由保真度、熵和成本感知的效用函数以及基于时序差分Q分数的epsilon-贪心规则驱动。我们在一个现实的城市物流测试平台上进行评估,该平台是一个基于德里地标、具有测地线节点间成本的真实CVRP,并在五种噪声族和四种严重程度上进行测试。在此实例上,QRL代理优于构造式启发式算法,并接近元启发式算法的性能,而误差缓解在高噪声下将近似比从0.84-0.87恢复到0.92-0.94。该控制器在短训练周期下从以CDR为主的机制转变为在较长训练周期下对四种技术的均衡部署,表明存在真正的机制依赖性选择。这些初步结果将自适应、学习感知的缓解定位为实现抗噪声QRL的实用途径。

英文摘要

Quantum Reinforcement Learning (QRL) represents policies as variational quantum circuits (VQCs), making it attractive for combinatorial optimization such as the Capacitated Vehicle Routing Problem (CVRP). On noisy intermediate-scale quantum (NISQ) hardware, however, decoherence degrades fidelity and destabilizes learning, and conventional error mitigation is applied statically without regard to the learning context. We introduce Adaptive Policy-Guided Error Mitigation (APGEM), a controller that selects among Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) online, driven by a fidelity, entropy, and cost aware utility function and an epsilon-greedy rule over temporal-difference Q-scores. We evaluate on a realistic urban-logistics testbed, a Delhi-based CVRP over real landmarks with geodesic inter-node costs, exercised across five noise families and four severity levels. On this instance, the QRL agent outperforms constructive heuristics and approaches metaheuristics, while mitigation restores approximation ratios from 0.84-0.87 to 0.92-0.94 under high noise. The controller shifts from a CDR-dominated regime under short training horizons to a balanced deployment across all four techniques under longer horizons, indicating genuine regime-dependent selection. These preliminary results position adaptive, learning-aware mitigation as a practical route to noise-resilient QRL.

补充信息

↑