发表机构
Base Labs(Base Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究训练5000万参数投影器将视觉编码器接入无原生视觉能力的大语言模型,探究仅扩大语言模型规模时视觉能力的扩展规律及可赋予的视觉能力边界。
AI 中文摘要
在冻结的视觉编码器与语言模型之间训练一个小型投影器,是一种成熟的多模态学习方法。随着语言模型参数数量的大幅增长,我们重新审视了在保持预训练权重不变的情况下,这种方法能够为模型增添哪些视觉能力。在此,我们训练了一个5000万参数的投影器,将Kimi K2.6的视觉编码器连接到GLM 5.2和5.3上,这两个模型本身不具备原生视觉能力,并进一步提出了一种可复现的、用于大规模训练这些适配器的方案。我们研究了以下问题:(a) 当仅扩大语言模型一侧的规模时,多模态模型的视觉能力如何随之扩展;(b) 在大规模条件下,哪些具体的视觉能力能够被赋予纯语言模型,而哪些仍然受限。我们在MMMU-Pro和BLINK上进行了评估,考察了整体性能以及各项具体视觉任务上的结果。
英文摘要
Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.
CommentsNeurIPS 2026 Workshop: Grounded and Faithful Vision-Language Models for Real-World Deployment