发表机构
School of Computing Technologies, RMIT University; Quantum Systems, Data61, CSIRO(RMIT大学计算技术学院; 澳大利亚联邦科学与工业研究组织数据61分部量子系统部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出DreamQAS模型,以基于模型的强化学习框架实现VQE高效量子架构搜索,在多个分子任务上降低了真实VQE调用次数并提升了反事实动作排名效用。
AI 中文摘要
基于强化学习的量子架构搜索(RL-QAS)在每次扩展量子电路后都会反复优化变分量子本征求解器(VQE),尽管电路构建和动作合法性是确定且已知的。我们提出DreamQAS,这是一种基于模型的强化学习框架,它保留了这些精确的电路动态,仅学习成本高昂的VQE后反馈。循环随机先验集成预测了相对于经验能量前沿的无预言机分数,并支持在显式合法电路上的多步想象策略学习。基于排名的激活、感知不确定性的悲观主义与截断,以及选择性真实VQE验证构成了一个可靠性控制的学习循环。在常见的15000回合预算和RL方法的冻结评估下,DreamQAS在5个分子任务中的4个上具有最低的平均冻结策略能量误差,在1个上为第二低。在两种方法的所有种子都达到的精细误差目标下,它在4个任务上使用的真实VQE调用减少了1.6倍至2.0倍,在BeH2-8q任务上减少了10.6倍。反事实动作排名效用在所有5个任务中均有所增加,平均增加0.346,95%置信区间为[0.185, 0.507],而直接贪婪和束搜索使用同一模型无法恢复想象策略学习的增益。集成分歧还在所有3个被研究任务上改善了风险覆盖,优于随机弃权(不执行)。这些结果确立了一种用于QAS的世界模型设计,其价值在于决策有用的反馈而非精确的能量预测。
英文摘要
Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly invokes a variational quantum eigensolver (VQE) after each gate addition even though circuit transitions and action legality are known. DreamQAS preserves these exact dynamics and learns only expensive post-VQE feedback through a recurrent ensemble that predicts a frontier-relative feedback score without requiring the exact ground-state energy, enabling uncertainty-controlled multi-step imagination. Under a common 15,000-episode budget and frozen evaluation, DreamQAS has the lowest reported mean error among RL methods on all five main molecular tasks. At fine-error targets reached by all seeds of DreamQAS and a matched non-imaginative control, it uses 1.6-2.0 times fewer real VQE calls on four tasks. Holding LiH-4q feedback-model weights fixed, its imagined-policy actor attains 0.073 mHa, versus 4.280 mHa and 4.434 mHa for greedy and beam deployment. Learned-transition and end-to-end predictor controls further show that preserving exact circuit structure and using feedback through policy learning are both important. Counterfactual action-ranking improves throughout training on all five probed tasks, while ensemble disagreement improves risk-coverage over random rejection on three tasks. DreamQAS therefore learns decision-useful feedback for QAS without modeling already-known circuit dynamics or requiring the exact ground-state energy.