arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合量子-经典NLP分类与紧凑语义表示:表示压缩的实验分析

Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression

Ali Hassan, Zijia Zhao, Maha A. Metawei

arXiv 2609.10089首次发表:更新:

发表机构

German University in Cairo; University of Melbourne; High Performance Computing Lab, Electronics Research Institute(开罗德国大学; 墨尔本大学; 高性能计算实验室,电子研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出混合量子-经典NLP分类流水线,比较PCA、NCA和LDA降维方法,发现监督降维在低维度下保留任务信息更有效,LDA和NCA在5维时接近经典基线性能。

AI 中文摘要

大型语言模型和句子嵌入模型提供了丰富的语义表示,但其高维度对近期量子机器学习(QML)构成挑战,因为量子电路只能处理有限数量的输入特征。我们研究了一种混合量子-经典流水线,将高维句子嵌入转换为用于变分量子分类的紧凑表示。该工作流程结合了预训练句子嵌入模型、降维、角度编码、变分量子电路(VQC)和经典决策层。我们系统比较了主成分分析(PCA)、邻域成分分析(NCA)和线性判别分析(LDA),涵盖无监督和有监督降维。使用TREC问题分类数据集,我们研究了表示维度、信息保留、量子比特数和分类性能之间的关系。初步PCA实验揭示了强烈的信息瓶颈:将768维嵌入降至3、4、5和8维时,分别保留约8.2%、10.2%、11.9%和16.4%的方差,相应的分类准确率为50.3%、51.2%、57.9%和63.4%。相比之下,有监督降维效率显著更高。在无泄漏交叉验证协议下,LDA仅使用5个维度达到85.3%的准确率,NCA达到83.1%,而完整的384维经典基线为85.1%。这些结果表明,有监督降维能比基于方差的压缩更有效地保留任务相关信息,使紧凑表示成为实现实用混合量子-经典NLP模型的有前景途径。

英文摘要

Large language and sentence-embedding models provide rich semantic representations, but their high dimensionality poses a challenge for near-term quantum machine learning (QML), where quantum circuits can process only a limited number of input features. We investigate a hybrid quantum-classical pipeline that transforms high-dimensional sentence embeddings into compact representations for variational quantum classification. The workflow combines a pretrained sentence-embedding model, dimensionality reduction, angle encoding, a variational quantum circuit (VQC), and a classical decision layer. We systematically compare principal component analysis (PCA), neighborhood components analysis (NCA), and linear discriminant analysis (LDA), covering both unsupervised and supervised dimensionality reduction. Using the TREC question-classification dataset, we study the relationship between representation dimensionality, information retention, qubit count, and classification performance. Preliminary PCA experiments reveal a strong information bottleneck: reducing 768-dimensional embeddings to 3, 4, 5, and 8 dimensions retains about 8.2%, 10.2%, 11.9%, and 16.4% of the variance, with corresponding classification accuracies of 50.3%, 51.2%, 57.9%, and 63.4%. In contrast, supervised reduction is substantially more efficient. LDA reaches 85.3% accuracy and NCA reaches 83.1% using only 5 dimensions, under a leakage-free cross-validation protocol, compared with 85.1% for a full 384-dimensional classical baseline. These results indicate that supervised dimensionality reduction can preserve task-relevant information far more effectively than variance-based compression, making compact representations a promising route toward practical hybrid quantum-classical NLP models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑