VIVAS:通过视觉-语言统一自回归监督激活视觉语言模型预训练中的视觉感知
VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision
- Tsinghua University(清华大学)
- Tencent Youtu Lab(腾讯优图实验室)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VIVAS通过统一标记空间和密集结构语义视觉分词器,在预训练中对视觉与语言内容进行统一自回归监督,以增强细粒度视觉感知,在12.4T标记上训练后于7项任务和39个多模态基准上达到最先进性能。
AI中文摘要:
尽管视觉语言模型(VLM)展现出强大的能力,但它们仍存在一个关键局限:细粒度视觉感知不足,这从根本上限制了其多模态理解能力。我们将这一瓶颈归因于预训练期间文本主导的优化偏差,这种偏差促使模型忽视细粒度的视觉细节,从而限制了多模态理解的能力。我们研究发现,克服这一瓶颈需要两个关键要素:(1)统一的标记空间范式,以确保稳定的训练动态;(2)一种模态对齐的密集视觉监督信号,该信号兼具结构粒度和语义信息,以捕获关键的视觉表征。基于这些见解,我们提出了VIVAS,这是一个建立在统一标记空间范式之上的框架,引入了一个密集-结构-语义视觉分词器,通过纳入视觉词汇将文本词汇扩展为统一的视觉-语言词汇。在预训练期间,VIVAS对视觉细节和语言内容执行视觉-语言统一自回归监督,从而增强视觉感知以改善多模态理解。VIVAS在12.4T个标记上进行端到端训练,在7项任务和39个多模态基准上取得了最先进的性能。
英文摘要:
While Vision-Language Models (VLMs) demonstrate strong capabilities, they continue to suffer from a critical limitation: insufficient fine-grained visual perception, which fundamentally limits their multimodal understanding. We attribute this bottleneck to text-dominant optimization biases during pre-training, which encourage the model to overlook fine-grained visual details, thereby limiting the capability of multimodal understanding. We investigate that overcoming this bottleneck requires two key elements: (1) a unified token space paradigm that ensures stable training dynamics, and (2) a modality-aligned dense visual supervision signal enriched with both structural granularity and semantic information to capture critical visual representations. Based on these insights, we propose VIVAS, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary. During pretraining, VIVAS performs vision-language unified autoregressive supervision over both visual details and linguistic content, thereby enhancing visual perception to improve multimodal understanding. Trained end-to-end on 12.4T tokens, VIVAS achieves state-of-the-art performance across 7 tasks and 39 multimodal benchmarks.