用强化学习建模量子神经网络梯度
Modeling quantum neural network gradient with reinforcement learning
浏览论文内容
中文总结 AI 辅助
提出 RLQ-Grad,用强化学习代理直接生成量子神经网络梯度,避免贫瘠高原和指数级微分成本,在多个基准上显著加速并提升准确率。
中文摘要 AI 辅助
在近期硬件上训练量子神经网络(QNN)仍受到两个叠加困难的阻碍:梯度方差指数级消失(即贫瘠高原),以及通过一个 $n$ 量子比特、$L$ 层电路进行微分所需的 $\mathcal{O}(L \cdot 2^n)$ 时间和内存成本。我们提出 RLQ-Grad,一种基于强化学习的优化器,其中经典策略 $\pi_\phi$(一个谱归一化的 PPO 智能体)学习直接提出参数更新,以 QNN 的当前参数、损失、准确率和先前的更新为条件。由于代理梯度由经典网络发出,而非通过对酉 $U(\theta)$ 进行微分获得,其方差不受贫瘠高原集中界限的约束,其成本随可训练参数数量而非希尔伯特空间维度扩展。我们正式证明了这些性质,并在一个硬件高效拟设上,在多达 $n=20$ 个量子比特的四个监督基准上进行了验证。RLQ-Grad 保持了近乎平坦的梯度方差曲线,而反向传播、参数平移和伴随微分则衰减 1 到 2 个数量级。考虑到完整训练流程(PPO rollout、actor-critic 更新和优化器状态),RLQ-Grad 在 $n=20$ 时内存占用低于 2 MB,每迭代运行速度分别比这三种方法快 $2490\times$、$7876\times$ 和 $673\times$。在多达 12 个量子比特的电路上,它相比基于梯度的基线将 top-1 准确率提升高达 $+10\\%$,并在 CIFAR-10 上 14 到 20 个量子比特处与专门的贫瘠高原缓解方法相匹配,而进化优化器和无梯度优化器则崩溃至随机水平。
英文摘要
Training quantum neural networks (QNNs) on near-term hardware remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the $\mathcal{O}(L \cdot 2^n)$ time and memory cost of differentiating through an $n$-qubit, $L$-layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy $π_ϕ$ (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary $U(θ)$, its variance is not constrained by the barren plateau concentration bound, and its cost scales with the number of trainable parameters rather than the Hilbert-space dimension. We prove these properties formally and verify them on a hardware-efficient ansatz across four supervised benchmarks with up to $n=20$ qubits. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1 to 2 orders of magnitude. Accounting for the full training pipeline (PPO rollouts, actor-critic updates, and optimizer states), RLQ-Grad needs under 2 MB of memory and runs $2490\times$, $7876\times$, and $673\times$ faster per iteration than these three methods at $n=20$. It improves top-1 accuracy by up to $+10\%$ over gradient-based baselines on circuits of up to 12 qubits, and matches dedicated barren plateau mitigation methods on CIFAR-10 at 14 to 20 qubits, where evolutionary and gradient-free optimizers collapse to chance.
发表机构
- College of Information and Communication Technology, Can Tho University(信息通信技术学院,芹苳大学)
- College of Engineering and Computer Science, VinUniversity(工程与计算机科学学院,VinUniversity)
- Center for Digital Transformation and Communications, Can Tho University(数字化转型与通信中心,芹苳大学)
- Department of Computer Science and Engineering, The University of Aizu(计算机科学与工程系,会津大学)
机构由 AI 辅助整理,请以论文原文为准。