arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13466cs.LGquant-ph

PQFA:用于多模态分类的融合表示的并行量子特征增强

PQFA: Parallel Quantum Feature Augmentation of Fused Representations for Multimodal Classification

Mingzhu Wang, Yun Shang

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态学习中后融合增强问题,提出并行量子特征增强框架PQFA,将其应用于融合特征,经多种处理后幅度编码到量子电路预测。实验表明PQFA性能优、参数效率高,是混合量子 - 经典多模态学习后融合增强的有效策略。

中文摘要 AI 辅助

大多数多模态学习方法致力于改进异构表示的对齐和融合,而后融合增强的探索较少。我们提出了并行量子特征增强(PQFA),这是一个混合量子 - 经典框架,将多个浅变分量子电路应用于融合的多模态特征。通过双向交叉注意力、注意力池化和自适应门控融合处理由冻结的RoBERTa和ViT编码器提取的文本和图像表示。然后将融合特征幅度编码到并行量子电路中,其测量读数与经典表示连接用于预测。我们在MM - IMDb和N24News上进行评估,通过控制比较使用相同的编码器、融合主干、数据分割、投影维度和增强输出宽度。PQFA始终优于没有量子增强的融合主干和宽度匹配的MLP增强基线,同时使用约2.2K增强参数,而MLP分支为24.0K。缺失模态实验表明,当文本或视觉输入不完整时,鲁棒性得到改善,当信息更丰富的文本模态严重退化时,增益尤为明显。控制消融和特征空间分析表明,随机特征映射、增加经典宽度或未训练的量子变换无法再现这种改进。量子态诊断还显示,在测试的模拟噪声水平和编码状态的不同分支特定变换中,预测性能稳定。这些结果表明PQFA是混合量子 - 经典多模态学习中后融合增强的有效且参数高效的策略。

英文摘要

Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features. Text and image representations extracted by frozen RoBERTa and ViT encoders are processed through bidirectional cross-attention, attentive pooling, and adaptive gated fusion. The fused feature is then amplitude-encoded into parallel quantum circuits, whose measurement readouts are concatenated with the classical representation for prediction. We evaluate PQFA on MM-IMDb and N24News through controlled comparisons using the same encoders, fusion backbone, data splits, projection dimension, and augmentation output width. PQFA consistently outperforms both the fusion backbone without quantum augmentation and a width-matched MLP augmentation baseline, while using approximately 2.2K augmentation parameters compared with 24.0K for the MLP branch. Missing-modality experiments further show improved robustness when textual or visual inputs are incomplete, with particularly clear gains when the more informative textual modality is severely degraded. Controlled ablations and feature-space analyses indicate that the improvement cannot be reproduced by random feature mappings, increased classical width, or untrained quantum transformations. Quantum-state diagnostics additionally show stable predictive performance across the tested simulated noise levels and distinct branch-specific transformations of the encoded states. These results establish PQFA as an effective and parameter-efficient strategy for post-fusion augmentation in hybrid quantum-classical multimodal learning.

发表机构

  • AMSS(中国科学院数学与系统科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑