发表机构
Anhui University; Origin Quantum Computing Company Limited(安徽大学; 本源量子计算有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型在少样本学习中细粒度区分不足的问题,提出多模态量子适配器MQAdapter。先检索Top-K候选类别作语义锚点,再用跨模态量子学习机制,利用量子特性建模高阶交互,细化视觉特征,参数高效且能集成现有算法提升性能。
AI 中文摘要
大规模视觉语言模型在广泛任务中展现出强大的迁移学习能力。在少样本分类中,视觉语言模型能有效筛选候选类别,Top-K准确率较高,但在视觉相似类别间的细粒度区分上存在困难,Top-1性能欠佳。现有的视觉语言模型适配器研究主要关注特征空间中视觉与文本表示的全局对齐,未利用语义相似类别细化细粒度视觉表示。基于此,我们提出了一种新颖的用于少样本学习的从粗到细的视觉语言模型微调方法——多模态量子适配器(MQAdapter)。具体而言,MQAdapter首先检索与输入图像最相似的Top-K类别候选者并将其用作语义锚点,然后采用跨模态量子学习机制在这些锚点的指导下细化视觉特征。该机制的核心是将视觉和文本特征编码为量子态。通过在高维希尔伯特空间中利用量子纠缠和叠加,MQAdapter有效地对高阶跨模态交互进行建模,产生比传统欧几里得适配器更具判别力的表示。MQAdapter参数高效,可与各种现有微调算法集成以进一步提升性能。对15个数据集的评估证明了MQAdapter的有效性,同时所需的可训练参数更少。
英文摘要
Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification, we observe that VLMs exhibit a notable ability to filter candidate categories and thus achieve high Top-K accuracy. However, they often struggle with fine-grained discrimination among visually similar categories, resulting in unsatisfactory Top-1 performance, as shown in Figure 1. Existing studies on VLM adapters generally focus on global alignment between visual and textual representations in the feature space, but fail to exploit semantically similar categories to refine fine-grained visual representations. Based on these observations, we propose a novel coarse-to-fine VLM fine-tuning approach for few-shot learning that leverages quantum computation, termed the Multi-Modal Quantum Adapter (MQAdapter). Specifically, MQAdapter first retrieves the Top-K category candidates most similar to the input image and uses them as semantic anchors. It then employs a cross-modal quantum learning mechanism to refine visual features under the guidance of these anchors. The core of this mechanism is the encoding of visual and textual features into quantum states. By leveraging quantum entanglement and superposition in a high-dimensional Hilbert space, MQAdapter effectively models higher-order cross-modal interactions, producing more discriminative representations than traditional Euclidean adapters. MQAdapter is parameter-efficient and can be integrated with various existing fine-tuning algorithms to achieve further performance gains. Evaluations on 15 datasets demonstrate the effectiveness of MQAdapter while requiring fewer trainable parameters.