发表机构
University College London; University of Amsterdam(伦敦大学学院; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态学习的组合概念泛化挑战,提出带多阶段训练的量子模型,在CLEVR数据集上验证其可提升分布外关系泛化能力且参数远少于经典基准。
AI 中文摘要
组合概念泛化(Compositional Concept Generalization,CoCoGen)是指在新场景中系统地重组已学习基元的能力,这是多模态学习的核心挑战。本研究提出一种意义的组合模型,将名词与关系分离,使用张量和变分量子电路在数据上进行训练。该模型支持多阶段训练范式:首先从单目标图像-文本对学习目标表征,随后将这些表征迁移到关系阶段,此时目标参数被冻结,仅优化关系组件。该设计在电路层面明确强制执行组合分解,确保关系作为稳定基元上的变换被学习。该训练范式在专为CoCoGen开发的CLEVR数据集上进行测试。对于文本,我们使用名词的向量表征和关系的高阶张量表征,采用一组不同的ansatz;对于图像,我们使用从OpenAI的视觉语言工具CLIP得到的图像嵌入的量子编码,以及保留原始嵌入几何的对比振幅编码,还有引入非线性特征变换的角度编码。结果表明,多阶段训练结合结构化编码显著提升了分布外关系泛化能力,同时可训练参数数量比经典基准少几个数量级。我们发现性能提升源于表征与编码的相互作用,非线性量子编码增强了组合结构的可分性。这些结果证明,结构化量子表征与阶段性学习为多模态量子机器学习中的组合泛化提供了有效框架。
英文摘要
Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning. In this work, we provide a solution using a compositional model of meaning that separates nouns from relations and uses tensors and variational quantum circuits to train them on data. This model enables us to employ a multi stage training paradigm, one that first learns object representations from single-object image-caption pairs, then subsequently transfers these to the relational stage where object parameters are frozen and optimisation is only applied to relational components. This design explicitly enforces compositional factorisation at the circuit, ensuring that relations are learned as transformations over stable primitives. The training paradigm is tested on the CLEVR dataset developed specificially for CoCoGen. For text, we work with vector representations of nouns and higher order tensor representations of relations using a set of different ansatz. For images, we work with quantum encodings of image embeddings dervied from Open AI's Vision Language tool CLIP and contrast amplitude encoding, which preserves the original embedding geometry, with angle encoding, which introduces nonlinear feature transformations. Our results show that multi-staged training combined with structured encodings significantly improves out of distribution relational generalisation, while using orders of magnitude fewer trainable parameters than classical baselines. We find that performance gains arise from the interaction between representation and encoding, with nonlinear quantum encodings enhancing the separability of compositional structure. These findings demonstrate that structured quantum representations and staged learning provide an effective framework for compositional generalisation in multimodal quantum machine learning.