MLLMCLIP:面向鲁棒视觉-语言表征的多模态大语言模型(MLLM)特征级蒸馏
MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations
浏览论文内容
中文总结 AI 辅助
MLLMCLIP是一种异构特征级蒸馏框架,直接将MLLM的多模态知识迁移至CLIP,在组合性任务和通用视觉-语言任务上均实现最优性能。
中文摘要 AI 辅助
CLIP等预训练视觉-语言模型在零样本识别中表现出色,但往往在组合性任务上失效,尤其是属性-对象和关系结构的建模。近期研究通过级联大语言模型与文本-图像模型生成合成难负例来增强训练,以此缓解该问题,但会产生大量流水线开销。我们提出MLLMCLIP,这是一种异构蒸馏框架,可直接将生成式多模态大语言模型(MLLM)教师的多模态知识迁移至判别式CLIP学生模型,完全绕过合成数据。为弥合两种范式的架构差异,我们引入基于注意力的逐层令牌选择机制和基于中心核对齐(CKA)的蒸馏损失函数。与现有CLIP增强方法相比,MLLMCLIP在组合准确率上达到当前最优水平,同时在标准零样本分类和图像-文本检索任务上实现持续提升,表明特征级蒸馏可同时增强组合性与通用视觉-语言表征能力。
英文摘要
Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting training with synthetic hard negatives generated by a cascade of large language models and text-to-image models, which incurs substantial pipeline overhead. We instead propose MLLMCLIP, a heterogeneous distillation framework that transfers multimodal knowledge directly from a generative Multimodal Large Language Model (MLLM) teacher into a discriminative CLIP student, bypassing synthetic data entirely. To bridge the architectural mismatch between the two paradigms, we introduce an attention-based per-layer token selection and a CKA-based distillation loss. Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capability.
发表机构
- KAIST(韩国科学技术院)
- Sony Group Corporation(索尼集团公司)
机构由 AI 辅助整理,请以论文原文为准。