arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35457cs.CV

我们距离移除视觉编码器还有多远?无编码器多模态预训练的缩放定律

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

  • CASIA UCAS Foundation Model Department, Tencent(中国科学院自动化研究所 中国科学院大学 腾讯基础模型部)

机构由 AI 辅助整理,请以论文原文为准。

Lin Chen, Bolin Ni, Qi Yang, Lan Jiang, Kun Ding, Xiaoran Fan, Hower Yang, Ying Wang, Shiming Xiang

AI总结:

本研究通过缩放定律比较,发现无编码器多模态大语言模型在足够计算量下可追平基于视觉编码器的模型,且视觉先验优势随规模减弱,表明无编码器架构是多模态预训练的有前景方向。

AI中文摘要:

大多数现代多模态大语言模型(MLLMs)都建立在预训练的视觉编码器之上,该编码器提供了强大的视觉先验。无编码器的MLLMs则直接从原始像素学习视觉表示,提供了一种简单且统一的架构,但其缩放行为尚未被系统地表征。为填补这一空白,我们比较了无编码器和基于编码器的MLLMs的缩放定律,并报告了三个主要发现:(1)移除视觉编码器将多模态目标的计算最优分配转向更大的模型,而对文本目标的计算最优分配几乎不变。(2)两种架构在文本目标上表现出几乎重叠的损失-计算前沿,但在多模态目标上出现分歧:无编码器模型在小规模下表现不佳,但预计在约10^22 FLOPs(浮点运算次数)处迎头赶上,这在实际预训练预算内完全可以达到。(3)在没有视觉编码器的情况下,语言模型通过视觉特定的适应来接管其角色:视觉token之间的双向交互随着训练计算的增加而变得日益有益,视觉处理转向更早的层,视觉token的专家路由变得更加集中。总体而言,我们的结果表明,预训练编码器提供的视觉先验的优势随规模增大而减弱,使无编码器架构成为多模态预训练的一个有前景的方向。

英文摘要:

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.

↑