arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估多模态大语言模型中视觉编码器的强基线

A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

Yilin Yang, Jun-Tao Tang, Kengyi Wang, Siyuan Su, Gaoyong Luo, Mingda Chen

arXiv 2610.05413首次发表:更新:

发表机构

Shanghai Jiao Tong University; Nanjing University; Fudan University(上海交通大学; 南京大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过大规模实验重新审视视觉编码器的跨模态评估,指出先前方法的局限性并提出简单改进,最终提出无训练的RAVEL方法,基于跨模态最近邻检索,以简单方式大幅超越先前方法,为MLLMs中的视觉编码器评估提供强基线。

AI 中文摘要

评估视觉编码器需要能够可靠预测其在多模态大语言模型(MLLMs)中下游性能的指标。尽管近期研究表明跨模态指标能更好地捕捉这种性能,但在实践中单模态指标仍是主导选择。在本工作中,我们通过大规模实验重新审视了视觉编码器的跨模态评估。我们识别出先前方法在实验设计和方法论表述上的重要局限性。在解决这些局限性并引入简单改进后,我们提出了RAVEL,一种基于跨模态最近邻检索的无训练方法。尽管其简单,RAVEL在我们的实验中达到了最先进的性能,大幅超越了先前方法。我们的结果表明,在仔细且全面的设置下评估的简单跨模态指标,可以为评估用于MLLMs的视觉编码器提供坚实基础。

英文摘要

Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.

CommentsCode is available at https://github.com/JuntaoTang/MLLM-VisionEncoder-Eval

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑