发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对量子器件调谐中参数串扰致学习不稳定问题,提出用在线学习动作空间分解表示的方法,框架QADAPT借此高效学习共享策略,实现零样本泛化及收敛步骤数恒定,为大规模量子处理器快速校准提供可扩展途径。
AI 中文摘要
合作多智能体强化学习适用于具有大参数空间和可利用局部结构的问题,如静电定义量子点阵列的调谐。但参数串扰强时,对单个智能体而言非平稳环境会破坏学习,手动调谐也受此困扰。我们提出在线学习动作空间的分解表示,以解耦智能体并最小化干扰。我们的框架QADAPT利用这种分解基于局部测量和奖励高效学习共享策略。通过此模块化策略,实现对未见量子器件尺寸的零样本泛化,并保持收敛到目标状态的步骤数大致恒定。这项工作为大规模量子处理器的快速校准提供了可扩展途径。
英文摘要
Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter cross-talk is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning - the same effect that plagues manual tuning of such systems. We propose using a factored representation of the action space, learned online, to decouple agents and minimize their interference. Our framework, QADAPT, uses this factorization to efficiently learn shared policies based on local measurements and rewards. With this modular strategy, we achieve zero-shot generalization to unseen quantum device sizes and maintain an approximately constant number of convergence steps to reach target regimes. This work provides a scalable route toward the rapid calibration of large-scale quantum processors.