arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于量子电子设计自动化(QEDA)相位组件中路由电荷奇偶项排序的屏蔽强化学习

Shielded RL for Route-Charged Parity-Term Ordering in QEDA Phase Components

Owen Friedewald, Ali Shiri Sichani, Chi-Ren Shyu

arXiv 2607.15307首次发表:更新:

AI 中文总结

研究QEDA相位组件中路由电荷奇偶项排序问题,采用屏蔽强化学习方法,以可行性屏蔽和精英策略结合特定代理训练,验证表明该方法对特定组件有效,能降低路由成本,不同电路有不同适用性。

AI 中文摘要

在量子电子设计自动化(QEDA)布局电路中,可交换相位项在重新排序下逻辑不变,但硬件映射后其路由成本因项顺序影响CNOT抵消、交互局部性和路由压力而大幅变化。我们将QEDA相位组件内的奇偶/支持相位项排序视为屏蔽强化学习问题:可行性屏蔽将每一步限制在未发射项,使每个轨迹都是有效排列,基于结合支持转移大小和重六边形拓扑距离特征的路由电荷代理训练精英(交叉熵方法)策略。通过将逻辑等效电路直接路由到合成IBM风格的重六边形映射进行验证。对于36项奇偶行走组件,学习到的排序将平均路由CX减少到336.0,在相同或更大代理预算下比2-opt和模拟退火搜索减少5.7 - 12.2%,比默认构建顺序减少22.3%;经邦费罗尼校正后,路由CX和路由深度增益显著。诚实转移审计表明代理对奇偶行走组件有预测性,但对提取量大或令牌/排列电路无预测性,贡献相应受限。

英文摘要

Commuting phase terms in quantum electronic design automation (QEDA) placement circuits are logically invariant under reordering, yet their routed cost varies substantially after hardware mapping, since term order affects CNOT cancellation, interaction locality, and routing pressure. We cast parity/support phase-term ordering within a QEDA phase component as a shielded reinforcement-learning problem: a feasibility shield restricts each step to unemitted terms, so every trajectory is a valid permutation by construction, and an elite (cross-entropy-method) policy is trained against a route-charged proxy combining support-transition size and heavy-hex topology-distance features. We validate by direct Qiskit routing of logically equivalent circuits to a synthetic IBM-style heavy-hex map. On 36-term parity-walk components (50 term seeds x 2 transpiler seeds, statistics at the term-seed level), the per-instance learned ordering reduces mean routed CX to 336.0, a 5.7-12.2% paired reduction over 2-opt and simulated-annealing search at equal or greater proxy budget and 22.3% over the default construction order; routed-CX and routed-depth gains are significant after Bonferroni correction. Honest transfer audits show the proxy is predictive for the parity-walk component but not for extraction-heavy or token/permutation circuits, which require architecture-aware rewards, scoping the contribution accordingly.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑