发表机构
The University of Chicago; University of Chinese Academy of Sciences(芝加哥大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ConvCue通过并行冻结CNN和门控交叉注意力增强VLM视觉特征,在两阶段训练下于13个基准上提升Qwen3-VL和LLaVA-OneVision的性能,平均分最高提升3.82。
AI 中文摘要
现代视觉语言模型(VLM)在广泛的多模态任务中表现出强大的性能,但在需要细粒度判别和空间理解的视觉问题上仍然存在困难。这些局限性促使我们研究是否可以在不替换其原生视觉编码器的情况下,通过补充视觉表示来改进现有的VLM。预训练的卷积网络提供了一种候选特征来源,其动机在于它们的局部连接性和空间权重共享。我们提出了CONVCUE,该方法通过一个并行的、冻结的预训练CNN的最终阶段特征来增强预训练VLM的原生视觉表示。一个可学习的适配器将卷积特征映射到原生视觉特征的维度,而门控交叉注意力允许原始视觉标记从CNN特征中检索信息。增强后的标记通过原始的视觉到语言投影器传递,并通过两阶段训练流程对模型进行适配。我们在Qwen3-VL-2B、Qwen3-VL-4B和LLaVA-OneVision-7B上评估了CONVCUE,覆盖了13个多模态基准,包括视觉问答、文档和图表理解以及多模态推理。CONVCUE在三个骨干网络上均提高了平均基准性能,超过了原始模型和匹配的两阶段微调对照。在Qwen3-VL-4B上,它在所有13个基准上均优于原始模型,并将平均得分从75.00提高到78.82(相对于匹配的微调对照)。这些结果表明,预训练的卷积表示通过学习的适配和融合集成后,可以在不替换原始视觉编码器的情况下提高现有VLM的视觉理解能力。
英文摘要
Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existing VLMs without replacing their native visual encoders. Pretrained convolutional networks offer a candidate feature source, motivated by their local connectivity and spatial weight sharing. We introduce CONVCUE, which augments the native visual representations of a pretrained VLM with final-stage features from a parallel, frozen pretrained CNN. A learnable adapter maps convolutional features to the native visual feature dimension, while gated cross-attention allows the original visual tokens to retrieve information from the CNN features. The enhanced tokens are passed through the original visual-to-language projector, and the model is adapted through a two-stage training procedure. We evaluate CONVCUE on Qwen3-VL-2B, Qwen3-VL-4B, and LLaVA-OneVision-7B across 13 multimodal benchmarks covering visual question answering, document and chart understanding, and multimodal reasoning. CONVCUE improves average benchmark performance over both the original models and matched two-stage fine-tuning controls on all three backbones. On Qwen3-VL-4B, it improves over the original model on all 13 benchmarks and raises the average score from 75.00 to 78.82 relative to the matched fine-tuning control. These results show that pretrained convolutional representations, when integrated through learned adaptation and fusion, can improve the visual understanding of existing VLMs without replacing their original visual encoders.