arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于QAOA的RL发现纠缠拓扑中的涌现问题-图对齐

Emergent Problem-Graph Alignment in RL-Discovered Entanglement Topologies for QAOA

Tobias Rohe, Federico Harjes Ruiloba, Markus Baumann, Gerhard Stenzel, Leo Sünkel, Thomas Gabor, Claudia Linnhoff-Popien

arXiv 2608.07686首次发表:更新:

AI 中文总结

该研究用RL智能体发现QAOA的纠缠拓扑,得到与问题图对齐的稀疏拓扑,在有限优化预算下性能优于全图拓扑,揭示了拓扑密度带来的可训练性与表达性的权衡。

AI 中文摘要

在量子近似优化算法(QAOA)中,量子比特对通过两量子比特门连接的纠缠拓扑,传统上被设置为与问题图的边集相等。这种耦合将电路设计与显式问题知识绑定,在有限的优化预算下可能无法产生最易训练的电路。我们研究强化学习(RL)智能体能否在不直接访问问题图的情况下,为基于QAOA的最大割优化发现更有效的纠缠拓扑。一个带掩码的近端策略优化(Masked Proximal Policy Optimization)智能体依次放置IsingZZ门以构建电路拓扑,而变分内层循环则优化得到的QAOA参数,并返回近似比作为稀疏终端奖励。智能体的观测仅包含目前已放置的边和当前近似比;图结构只能通过优化奖励间接推断。在最多10个量子比特的Erdős–Rényi实例上,尽管智能体的观测中没有收到关于图结构的显式信息,它仍一致收敛到问题图的严格子集拓扑,实现了接近1.0的重叠率。在优化预算有限(50个梯度步骤)时,这些稀疏的、与问题对齐的拓扑优于全图拓扑和若干结构基线,但在优化预算充足时,会被更密集的拓扑超越。我们的结果揭示了由拓扑密度决定的可训练性-表达性权衡,并表明变分优化地形隐含编码了问题哈密顿量的结构信息。

英文摘要

In the Quantum Approximate Optimization Algorithm (QAOA), the entanglement topology, where qubit pairs are connected by two-qubit gates, is conventionally set equal to the edge set of the problem graph. This coupling ties circuit design to explicit problem knowledge and may not yield the most trainable circuit under limited optimization budgets. We investigate whether a reinforcement learning (RL) agent can discover more effective entanglement topologies for QAOA-based MaxCut optimization without direct access to the problem graph. A Masked Proximal Policy Optimization agent sequentially places IsingZZ gates to construct a circuit topology, while a variational inner loop optimizes the resulting QAOA parameters and returns the approximation ratio as a sparse terminal reward. The agent's observation contains only the edges placed so far and the current approximation ratio; graph structure can only be inferred indirectly through the optimization reward. On Erdős--Rényi instances with up to $10$~qubits, the agent consistently converges to topologies that are strict subsets of the problem graph, achieving overlap ratios approaching $1.0$, despite receiving no explicit information about the graph structure in its observations. These sparse, problem-aligned topologies outperform the full graph topology and several structural baselines when the optimization budget is limited ($50$~gradient steps), but are overtaken by denser topologies given sufficient optimization budget. Our results reveal a trainability--expressibility trade-off governed by topology density and suggest that the variational optimization landscape implicitly encodes structural information about the problem Hamiltonian.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑