arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27948cs.CV

VIVAS:通过视觉-语言统一自回归监督激活视觉语言模型预训练中的视觉感知

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

  • Tsinghua University(清华大学)
  • Tencent Youtu Lab(腾讯优图实验室)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun

AI总结:

VIVAS通过统一标记空间和密集结构语义视觉分词器,在预训练中对视觉与语言内容进行统一自回归监督,以增强细粒度视觉感知,在12.4T标记上训练后于7项任务和39个多模态基准上达到最先进性能。

AI中文摘要:

尽管视觉语言模型(VLM)展现出强大的能力,但它们仍存在一个关键局限:细粒度视觉感知不足,这从根本上限制了其多模态理解能力。我们将这一瓶颈归因于预训练期间文本主导的优化偏差,这种偏差促使模型忽视细粒度的视觉细节,从而限制了多模态理解的能力。我们研究发现,克服这一瓶颈需要两个关键要素:(1)统一的标记空间范式,以确保稳定的训练动态;(2)一种模态对齐的密集视觉监督信号,该信号兼具结构粒度和语义信息,以捕获关键的视觉表征。基于这些见解,我们提出了VIVAS,这是一个建立在统一标记空间范式之上的框架,引入了一个密集-结构-语义视觉分词器,通过纳入视觉词汇将文本词汇扩展为统一的视觉-语言词汇。在预训练期间,VIVAS对视觉细节和语言内容执行视觉-语言统一自回归监督,从而增强视觉感知以改善多模态理解。VIVAS在12.4T个标记上进行端到端训练,在7项任务和39个多模态基准上取得了最先进的性能。

英文摘要:

While Vision-Language Models (VLMs) demonstrate strong capabilities, they continue to suffer from a critical limitation: insufficient fine-grained visual perception, which fundamentally limits their multimodal understanding. We attribute this bottleneck to text-dominant optimization biases during pre-training, which encourage the model to overlook fine-grained visual details, thereby limiting the capability of multimodal understanding. We investigate that overcoming this bottleneck requires two key elements: (1) a unified token space paradigm that ensures stable training dynamics, and (2) a modality-aligned dense visual supervision signal enriched with both structural granularity and semantic information to capture critical visual representations. Based on these insights, we propose VIVAS, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary. During pretraining, VIVAS performs vision-language unified autoregressive supervision over both visual details and linguistic content, thereby enhancing visual perception to improve multimodal understanding. Trained end-to-end on 12.4T tokens, VIVAS achieves state-of-the-art performance across 7 tasks and 39 multimodal benchmarks.

↑