arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.21447cs.CVcs.AI

多模态语言模型在更浅层视觉特征下看得更清楚

Multimodal Language Models See Better When They Look Shallower

  • Zhejiang Gongshang University(浙江工商大学)
  • Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
  • Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波)
  • Meituan Inc.(美团公司)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

更新

AI总结:

本研究首次系统探究多模态大语言模型的视觉层选择,发现浅层特征在细粒度视觉任务上优于深层,并提出轻量级融合方法提升性能。

AI中文摘要:

多模态大语言模型(MLLMs)通常从预训练的视觉Transformer(ViT)的最终层提取视觉特征。这种普遍的深层偏好主要源于经验惯例,而非基于原理的分析。尽管先前的研究表明,不同的ViT层捕获不同类型的信息,较浅层侧重于精细的视觉细节,而较深层更贴近文本语义,但这种差异对MLLM性能的影响仍未得到充分探索。我们首次对MLLM的视觉层选择进行了全面研究,分析了ViT各层之间的表示相似性,以建立浅层、中层和深层分组。通过对参数规模在1.4B至7B之间的MLLM在涵盖60多个任务的10个基准上进行广泛评估,我们发现,虽然深层在OCR等语义丰富的任务中表现出色,但浅层和中间层在计数、定位和物体定位等细粒度视觉任务上显著优于深层。基于这些见解,我们提出了一种轻量级特征融合方法,策略性地整合较浅层,在单层和专门融合基线上均取得了一致的改进。我们的工作首次对MLLM中的视觉层选择进行了原理性研究,表明MLLM在查看更浅层时往往能看得更清楚。

英文摘要:

Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT). This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis. While prior studies suggest that different ViT layers capture different types of information, with shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored. We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings. Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization. Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines. Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower.

补充信息

↑