“训练阶段用经典计算,部署阶段用量子计算”需要重新思考泛化性
"Train classical, deploy quantum" requires rethinking generalization
浏览论文内容
中文总结 AI 辅助
研究针对“训练用经典、部署用量子”的量子生成模型范式,发现矩匹配损失训练的模型泛化性弱于似然训练模型,需直接针对泛化性设计方法。
中文摘要 AI 辅助
生成模型已成为科学与工业领域的核心工具,应用涵盖图像与文本合成、分子及材料设计等。量子生成模型被视为量子计算机最具前景的应用之一,因为量子电路可自然生成其编码分布的样本,且对于合适的电路,该分布被认为难以被任何经典计算机复现。主流策略是在经典计算机上训练这些模型,部署阶段则使用量子设备生成样本,当训练损失可在经典计算机上评估时该策略可行,典型例子是最大均值差异(MMD²),这是一种矩匹配损失,通过模型与数据的泡利Z相关性对两者进行比较。迄今为止的研究已探讨此类模型能否训练、其采样是否困难,但最小化该目标是否会产生具有泛化性的模型,而非仅复现训练统计量,这一问题仍未得到充分理解。我们通过直接采样对大量量子与经典生成模型进行基准测试,发现采用矩匹配损失训练的模型通常比似然训练的模型泛化性更差,这一结论在两个受应用启发的数据集上得到验证:第一个是最多30量子比特的基数约束数据集,第二个是基因组单核苷酸变异数据集,其有效集为观测数据。这些结果表明,收敛的矩匹配损失并非泛化性的可靠衡量标准,“训练阶段用经典计算,部署阶段用量子计算”的工作流将需要直接针对泛化性的方法,而更好的训练目标是否足够或模型架构本身必须改变仍未可知。
英文摘要
Generative models have become central across science and industry, from image and text synthesis to the design of molecules and materials. Quantum generative models are considered one of the most promising applications for quantum computers, since a quantum circuit naturally produces samples from the distribution it encodes, and for suitable circuits that distribution is believed to be hard for any classical computer to reproduce. A leading strategy trains these models on a classical computer and reserves the quantum device for generating samples at deployment. This is possible when the training loss can be evaluated on a classical computer. A prime example is the maximum mean discrepancy (MMD$^2$), a moment-matching loss that compares the model and the data through their Pauli-$Z$ correlations. Research so far has asked whether such models can be trained and whether their sampling is hard; whether minimizing such an objective yields a model that \emph{generalizes}, rather than one that merely reproduces the training statistics, remains poorly understood. We benchmark thirteen quantum and classical generative models by direct sampling on two application-inspired datasets: first a cardinality-constrained dataset at up to $30$ qubits and second a dataset of genomic single-nucleotide variants, whose valid set is the observed data. Models that converge the loss to the same value differ widely in how much of the unseen valid set they cover. These results indicate that a converged moment-matching loss is not a reliable measure of generalization, and that a train-classical, deploy-quantum workflow has to measure generalization by sampling the trained model, a step that at the sizes of interest is believed to require the quantum device.
发表机构
- QC Ware Corp.(QC Ware公司)
- Sorbonne Université(索邦大学)
机构由 AI 辅助整理,请以论文原文为准。