量子特征工程用于信用违约预测:IQP电路何时以及为何帮助线性分类器
Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers
浏览论文内容
中文总结 AI 辅助
本研究使用IQP电路生成特征,在信用违约预测中提升逻辑回归F1分数0.055,优于核主成分分析,且特征选择至关重要,揭示了线性表达能力机制。
中文摘要 AI 辅助
信用违约预测是一个表格分类问题,其中F1分数的适度提升可直接转化为金融风险的降低。我们探究瞬时量子多项式时间(IQP)电路能否产生特征,在相同特征预算下,使分类器优于其原始经典基线和核主成分分析(Kernel PCA)——最强的无监督经典非线性替代方法。数据集为每位客户提供23个金融属性;对于n量子比特电路,我们从中选取n个属性,将每个属性编码为旋转角度,并读出2n个期望值作为新特征。使用量子电路的动机在于计算层面:n量子比特的IQP电路以恒定深度运行,并在2^n维希尔伯特空间中编码特征相关性,而对其精确输出统计的经典模拟随n呈指数级增长。使用UCI信用卡客户违约数据集和五折交叉验证,我们发现将16个IQP特征(n=8量子比特)附加到逻辑回归模型上,F1分数从0.462提升至0.517(+0.055,p<0.0001)。核主成分分析作为次优方法,在相同特征数量下仅达到0.493;该差距在12项测试的Benjamini-Hochberg校正后依然显著(p=0.00007)。其他分类器——随机森林、支持向量机、XGBoost或k近邻——均未受益,这表明存在线性表达能力机制而非通用改进。我们还展示了8个输入特征的选择方式至关重要:基于随机森林重要性的选择达到F1=0.523,而编码最大不相关特征则降至0.496,证明电路放大信息结构而非凭空创造。
英文摘要
Credit default prediction is a tabular classification problem in which modest gains in F1 translate directly into reduced financial exposure. We ask whether Instantaneous Quantum Polynomial-time (IQP) circuits can produce features that improve a classifier over both its raw classical baseline and Kernel PCA - the strongest unsupervised classical non-linear alternative - at an equal feature budget. The dataset provides 23 financial attributes per client; for an n-qubit circuit we select n of them, encode each as a rotation angle, and read 2n expectation values back out as new features. The motivation for using a quantum circuit is computational: an n-qubit IQP circuit runs in constant depth and encodes feature correlations in a 2^n-dimensional Hilbert space, whereas classical simulation of its exact output statistics scales exponentially in n. Using the UCI Default of Credit Card Clients dataset and five-fold cross-validation, we find that appending 16 IQP features (n = 8 qubits) to a Logistic Regression model raises F1 from 0.462 to 0.517 (+0.055, p < 0.0001). Kernel PCA, the next-best method, reaches only 0.493 at the same feature count; the gap survives Benjamini-Hochberg correction across 12 tests (p = 0.00007). No other classifier - Random Forest, SVM, XGBoost, or k-NN - benefits, which points to a linear-expressivity mechanism rather than a generic improvement. We also show that how the 8 input features are chosen matters: Random Forest importance-guided selection reaches F1 = 0.523, while encoding maximally uncorrelated features drops it to 0.496, demonstrating that the circuit amplifies informative structure rather than creating it from scratch.
发表机构
- Reichman University(赖希曼大学)
机构由 AI 辅助整理,请以论文原文为准。