面向化学基础模型的量子机器学习的离散化感知微调
Discretization-Aware Fine-Tuning for Quantum Machine Learning with Chemical Foundation Models
浏览论文内容
中文总结 AI 辅助
本研究提出离散化感知微调(DAFT)方法,在信息受限场景下使量子模型通过对齐连续表示与离散量子编码,在10量子比特的血脑屏障穿透预测任务中实现了对经典基线的量子优势。
中文摘要 AI 辅助
实际量子机器学习(QML)面临的一个关键挑战,尤其在分类等判别任务中,是近期量子设备将高维经典数据编码到小型量子寄存器的能力有限。在优化的基编码(比特-比特)设置中,这一约束会导致跨类碰撞,即不同标签的样本被映射到相同的离散比特串,从而对任何下游模型都不可区分。本研究探讨了在这种严重信息瓶颈下数据表示如何影响QML性能,提出了离散化感知微调(DAFT)方法,该方法使预训练的化学基础模型生成在量化后仍保持信息性的表示,通过可微分软碰撞损失降低碰撞概率。在受控设置下,我们评估了量子模型和经典模型,它们接收相同的离散比特串输入,从而将表示的影响与模型架构的影响隔离开来。在使用ChemBERTa-77M的血脑屏障穿透(BBBP)分子属性预测基准上,与冻结主干相比,DAFT将碰撞数量降低了数个数量级,并将量子分类准确率提高了12个百分点以上。重要的是,没有DAFT时,在相同输入约束下经典模型优于QML;而使用DAFT后,在更高量子比特数下这种比较发生反转:在10个量子比特时,量子模型超过了在相同比特串上训练的匹配经典基线(0.883对0.855,p=0.026)。这些结果表明,在信息受限的场景中,实现量子优势关键在于使连续表示与离散量子编码对齐。
英文摘要
A key challenge in practical quantum machine learning (QML), particularly for discriminative tasks such as classification, is the limited capacity of near-term quantum devices to encode high-dimensional classical data into small quantum registers. In optimized basis-encoded (bit-bit) settings, this constraint leads to cross-class collisions, where samples with different labels are mapped to the same discrete bit-string and thus become indistinguishable to any downstream model. In this work, we investigate how data representation affects QML performance under such severe information bottlenecks. We introduce discretization-aware fine-tuning (DAFT), a method that adapts a pre-trained chemical foundation model to produce representations that remain informative after quantization. DAFT reduces collision probability through a differentiable soft collision loss. We evaluate both quantum and classical models under a controlled setting in which they receive identical discretized bit-string inputs, isolating the effect of representation from model architecture. On the blood-brain barrier penetration (BBBP) molecular property prediction benchmark using ChemBERTa-77M, DAFT reduces collision counts by several orders of magnitude and improves quantum classification accuracy by more than 12 percentage points compared to a frozen backbone. Importantly, without DAFT, classical models outperform QML under the same input constraints. With DAFT, however, this comparison reverses at higher qubit counts. At 10 qubits, the quantum model surpasses a matched classical baseline trained on identical bit-strings (0.883 vs. 0.855, $p = 0.026$). These results show that, in information-constrained regimes, achieving a quantum advantage critically depends on aligning continuous representations with discrete quantum encodings.
发表机构
- RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences (iTHEMS), RIKEN(RIKEN跨学科理论与数学科学中心(iTHEMS))
- University of British Columbia(不列颠哥伦比亚大学)
- Japan Agency for Marine-Earth Science and Technology(日本海洋地球科学技术局)
- University of Guelph(圭尔夫大学)
- Cascade Quantum
机构由 AI 辅助整理,请以论文原文为准。