发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可分离定律,基于26个VLM模型在高分辨率基准上的测量,预测规模化对性能的影响,并为骨干网络大小与视觉标记数量间的计算分配提供闭式规则。
AI 中文摘要
视觉-语言模型(VLMs)在固定预算下面临权衡:处理更多视觉信息以实现细粒度感知,或使用更大的语言骨干网络进行复杂推理。现有研究并未告诉我们,在给定部署场景下,应选择何种骨干网络大小与输入分辨率的组合,尤其是在高分辨率部署中。为填补这一空白,我们提出了“可分离定律”,该定律描述了VLM性能如何随语言骨干网络大小和视觉标记数量变化。我们基于26个InternVL和QwenVL模型(语言骨干网络规模从1B到72B)的测量结果,在四个高分辨率基准(图像尺寸从224像素到8K)上拟合该定律。我们发现,能够响应规模化的查询可根据其所需技能进行预测,而相当一部分查询则完全不受规模化影响。我们还发现,两个模型系列在增大骨干网络时获益相似,但在增加视觉标记时获益差异显著。结合成本定律,可分离定律为在骨干网络大小和视觉标记之间分配计算资源提供了闭式规则。当部署受限于可用配置时,该定律可识别出在相同预算下接近最优可行选择的模型和图像尺寸。我们希望这项工作能为在高分辨率下根据模型需推理的内容决定其可见范围提供一种原则性方法。
英文摘要
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.