arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00013cs.CLcs.AIcs.CV

从文本到视觉的可迁移性:视觉语言模型(VLM)的能力缩放定律与迁移动态

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

  • Meituan(美团)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai

AI总结:

该研究提出首个跨家族的能力驱动多模态缩放定律,可通过LLM文本能力预测VLM性能,将主干选择从经验搜索转为定量决策,还揭示了LLM作为VLM主干的相关见解。

AI中文摘要:

构建视觉语言模型(VLM)时,选择合适的大语言模型(LLM)主干是最关键的决策,但目前仍缺乏原则性方法:基于计算的缩放定律无法跨模型家族推广,且尚无框架能在训练开始前直接预测VLM性能。我们提出能力驱动的多模态缩放定律,这是首个可通过直接观测的文本能力预测VLM基准准确率的跨家族框架。给定通过主成分分析(PCA)从LLM文本基准中提取的低维能力分数S,我们将VLM性能建模为S的函数,其中包含每个主干的迁移率和量化数据缩放效率的吸收率。为拟合和验证该框架,我们在严格控制的配置下,针对7个模型家族的34个LLM训练了超过150个VLM。对200多个文本基准和50多个多模态基准的评估表明,该定律能准确将迁移率从参数达80亿的模型外推至720亿规模的主干,高精度预测完整的VLM训练轨迹,并推广至完全未见过的模型家族。除缩放定律外,我们的分析还揭示了可操作的见解:某些文本基准与多模态性能负相关,暴露了潜在的基准博弈行为;基础LLM作为VLM主干优于指令微调的对应模型,因其吸收率更高且数据缩放衰减更低;不同模型家族在迁移-吸收率空间中占据不同位置。该框架将主干选择从成本高昂的经验搜索转变为有原则的定量决策。代码和数据可在https URL获取。

英文摘要:

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at https://github.com/wangq-dev/CDMScaling.

相关深度报道

↑