arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRPO-QM:用于量子层析的目标保持探索

GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling

Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling

arXiv 2609.14711首次发表:更新:

发表机构

Stony Brook University; Georgia Institute of Technology; Westlake University(石溪大学; 佐治亚理工学院; 西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GRPO-QM通过仅学习量子层析后验的探索策略并采用Metropolis校正保持后验平稳,实验表明大部分增益来自物理提议和先验而非学习,并诊断出采样训练中的目标缩放问题。

AI 中文摘要

基于奖励的学习可能会改变科学推断旨在估计的后验分布本身。GRPO-QM通过仅针对给定的量子层析后验学习探索策略来规避这一问题:群体相对策略在可逆的物理移动中进行选择,并且精确的Metropolis校正确保策略固定后后验保持平稳。然后我们考察学习在物理提议机制和先验知识之外贡献了什么。重建比较表明,相对于测试的流模型,大部分增益来自这两个组成部分,而非学习本身;一个闭式反例解释了原因:与接受移动相关的奖励即使在物理可观测量保持高度相关时也可能增加。控制初始化的奖励比较还表明,跨后验聚合可能会颠倒奖励在单个后验内隐含的排序。最后,在45个枚举的后验上,每个后验使用三个训练种子,相同轨迹目标的精确和采样梯度相对于调优的混合模型分别将平均物理估计方差降低了$12.63\%$和$6.32\%$,而不重新缩放其惩罚的轨迹得分平均仅降低$0.43\%$。这表明目标缩放解释了采样训练与精确训练之间差异的一部分。恢复的收益集中在四次测量时,并在十六次测量时符号翻转,因此这些精确的、可见的库诊断识别了采样训练中一个具体且可复现的失败模式、解决部分问题的目标缩放修复以及剩余的差距。

英文摘要

Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training. Across the evaluated reconstruction benchmarks, GRPO-QPS improves over BuresTomFlow and Flow-GRPO on thermal, cat, Dicke, and cluster families, and it closely matches an exact two-qubit reference posterior. Tuned conventional MCMC is slightly stronger on several original continuous benchmarks where the available fixed proposals already match the posterior geometry well. To test whether this reflects a fundamental limitation of learned exploration, we evaluate a more challenging multimodal thermal posterior. At six qubits and 800 shots, the learned proposal achieves a minimum effective sample size of 102 per $1{,}000$ likelihood calls, compared with 28 for prior independence, 27 for a tuned fixed mixture, and 20 for Haario adaptive Metropolis. A record-conditioned policy also transfers to unseen 3,000-shot records, matching or exceeding the strongest conventional baseline in all nine held-out seed-record comparisons. These results show that GRPO-QPS combines target-preserving Bayesian inference with broad gains over learned transport baselines and a sampling advantage when efficient exploration requires proposal geometry beyond the evaluated conventional kernels.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑