量子纠缠多模态融合网络(QEMFN):通过可训练纠缠实现资源感知的混合视觉-语言融合
Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement
浏览论文内容
中文总结 AI 辅助
提出QEMFN混合量子-经典框架,用可训练纠缠作为多模态融合的归纳偏置,在COCO-5k和Flickr30k上超越经典基线,并进行了受控消融和量子中心分析。
中文摘要 AI 辅助
多模态视觉-语言系统通常通过经典算子(如拼接、注意力、双线性池化或张量交互)融合图像和文本嵌入。我们提出量子纠缠多模态融合网络(QEMFN),这是一个混合量子-经典框架,将参数化纠缠作为结构化归纳偏置引入多模态融合。预训练的视觉和文本特征被投影到紧凑的潜在空间,编码为角度参数化的量子态,通过模态内和配对跨模态纠缠电路处理,并测量以产生用于检索的融合表示。在匹配的参数预算和相同的冻结CLIP骨干下,QEMFN在COCO-5k和Flickr30k上优于经典融合基线,包括多层感知机、张量融合、FiLM、交叉注意力、紧凑Transformer以及去量子化的配对拓扑模拟。一个消融套件将量子模块的贡献与周围的经典投影分离,量子中心分析报告了Meyer-Wallach纠缠能力、可表达性、梯度方差(对照贫瘠高原界限)以及控制训练进度下的熵-性能相关性,同时进行了对纠缠组件的干预研究。QEMFN在基于采样的估计、噪声模拟的假后端以及带有零噪声外推的真实超导设备上执行。这项工作不声称量子计算优势;贡献在于框架本身,以及受控的经验和量子中心评估,将可训练纠缠定位为在当代设备可访问规模上的可解释、可硬件执行的融合机制。
英文摘要
Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.
发表机构
- University of Missouri(密苏里大学)
机构由 AI 辅助整理,请以论文原文为准。