arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08216cs.AIquant-ph

量子纠缠多模态融合网络(QEMFN):通过可训练纠缠实现资源感知的混合视觉-语言融合

Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement

Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu

首次发表
浏览论文内容

中文总结 AI 辅助

提出QEMFN混合量子-经典框架,用可训练纠缠作为多模态融合的归纳偏置,在COCO-5k和Flickr30k上超越经典基线,并进行了受控消融和量子中心分析。

中文摘要 AI 辅助

多模态视觉-语言系统通常通过经典算子(如拼接、注意力、双线性池化或张量交互)融合图像和文本嵌入。我们提出量子纠缠多模态融合网络(QEMFN),这是一个混合量子-经典框架,将参数化纠缠作为结构化归纳偏置引入多模态融合。预训练的视觉和文本特征被投影到紧凑的潜在空间,编码为角度参数化的量子态,通过模态内和配对跨模态纠缠电路处理,并测量以产生用于检索的融合表示。在匹配的参数预算和相同的冻结CLIP骨干下,QEMFN在COCO-5k和Flickr30k上优于经典融合基线,包括多层感知机、张量融合、FiLM、交叉注意力、紧凑Transformer以及去量子化的配对拓扑模拟。一个消融套件将量子模块的贡献与周围的经典投影分离,量子中心分析报告了Meyer-Wallach纠缠能力、可表达性、梯度方差(对照贫瘠高原界限)以及控制训练进度下的熵-性能相关性,同时进行了对纠缠组件的干预研究。QEMFN在基于采样的估计、噪声模拟的假后端以及带有零噪声外推的真实超导设备上执行。这项工作不声称量子计算优势;贡献在于框架本身,以及受控的经验和量子中心评估,将可训练纠缠定位为在当代设备可访问规模上的可解释、可硬件执行的融合机制。

英文摘要

Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.

发表机构

  • University of Missouri(密苏里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑