发表机构
Department of Computer Engineering King Mongkut's University of Technology Thonburi Bangkok, Thailand
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器学习中数据稀缺和类不平衡问题,提出用QCBM的混合量子 - 经典框架生成合成数据。利用量子特性建模复杂概率分布,经实验验证其在扩充训练数据、提升分类性能及结构相似性上效果良好,可作为数据增强的补充工具用于不平衡表格数据。
AI 中文摘要
数据稀缺和类不平衡是机器学习中持续存在的挑战,会降低模型泛化能力并引入预测偏差。我们提出了一个使用量子电路玻恩机器(QCBM)的混合量子 - 经典合成数据生成框架来解决这些限制。该方法利用参数化变分量子电路中的量子力学特性——叠加和纠缠——来建模经典生成方法难以捕捉的复杂概率分布。在鸢尾花数据集和电信客户流失数据集上进行了实验。预处理包括归一化和基于主成分分析的降维,以实现量子电路的高效基编码。通过基于梯度的参数移位优化规则最小化真实数据和生成数据分布之间的库尔贝克 - 莱布勒(KL)散度来训练QCBM。用QCBM生成的合成样本以少数类的40 - 50% 扩充训练数据,可使F1分数提高约5 - 15%,少数类召回率提高10 - 25%。跨域评估显示性能差距仅为3 - 10%,表明分布保真度高。与经典过采样方法的比较分析表明,QCBM在电信数据集上实现了有竞争力的分类性能,并产生了更低的最大均值差异(MMD),这表明在某些不平衡设置中具有优越的结构相似性。这些发现确立了QCBM作为数据增强的可行补充工具,特别是对于具有类不平衡的低维结构化表格数据。
英文摘要
Data scarcity and class imbalance are persistent challenges in machine learning that degrade model generalization and introduce predictive bias. We present a hybrid quantum-classical framework for synthetic data generation using a Quantum Circuit Born Machine (QCBM) to address these limitations. The proposed approach exploits quantum mechanical properties -- superposition and entanglement -- within a parameterized variational quantum circuit to model complex probability distributions that are difficult for classical generative methods to capture. Experiments are conducted on two tabular benchmark datasets: the Iris dataset and the Telco Customer Churn dataset. Preprocessing includes normalization and PCA-based dimensionality reduction to enable efficient basis encoding for quantum circuits. The QCBM is trained by minimizing Kullback-Leibler (KL) divergence between real and generated data distributions using a gradient-based parameter-shift optimization rule. Augmenting training data with QCBM-generated synthetic samples at 40-50% of the minority class improves F1-score by approximately 5-15% and minority-class recall by 10-25%. Cross-domain evaluations (Train on Synthetic, Test on Real; and Train on Real, Test on Synthetic) reveal a performance gap of only 3-10%, indicating strong distributional fidelity. Comparative analysis against classical oversampling methods -- SMOTE, Borderline-SMOTE, KMeansSMOTE, and SVM-SMOTE -- shows that QCBM achieves competitive classification performance and produces lower Maximum Mean Discrepancy (MMD) on the Telco dataset, suggesting superior structural similarity in certain imbalanced settings. These findings establish QCBM as a viable complementary tool for data augmentation, particularly for low-dimensional structured tabular data with class imbalance.