arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25575cs.CVcs.AI

MLLMCLIP:面向鲁棒视觉-语言表征的多模态大语言模型(MLLM)特征级蒸馏

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji

首次发表
浏览论文内容

中文总结 AI 辅助

MLLMCLIP是一种异构特征级蒸馏框架,直接将MLLM的多模态知识迁移至CLIP,在组合性任务和通用视觉-语言任务上均实现最优性能。

中文摘要 AI 辅助

CLIP等预训练视觉-语言模型在零样本识别中表现出色,但往往在组合性任务上失效,尤其是属性-对象和关系结构的建模。近期研究通过级联大语言模型与文本-图像模型生成合成难负例来增强训练,以此缓解该问题,但会产生大量流水线开销。我们提出MLLMCLIP,这是一种异构蒸馏框架,可直接将生成式多模态大语言模型(MLLM)教师的多模态知识迁移至判别式CLIP学生模型,完全绕过合成数据。为弥合两种范式的架构差异,我们引入基于注意力的逐层令牌选择机制和基于中心核对齐(CKA)的蒸馏损失函数。与现有CLIP增强方法相比,MLLMCLIP在组合准确率上达到当前最优水平,同时在标准零样本分类和图像-文本检索任务上实现持续提升,表明特征级蒸馏可同时增强组合性与通用视觉-语言表征能力。

英文摘要

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting training with synthetic hard negatives generated by a cascade of large language models and text-to-image models, which incurs substantial pipeline overhead. We instead propose MLLMCLIP, a heterogeneous distillation framework that transfers multimodal knowledge directly from a generative Multimodal Large Language Model (MLLM) teacher into a discriminative CLIP student, bypassing synthetic data entirely. To bridge the architectural mismatch between the two paradigms, we introduce an attention-based per-layer token selection and a CKA-based distillation loss. Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capability.

发表机构

  • KAIST(韩国科学技术院)
  • Sony Group Corporation(索尼集团公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑