arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27915cs.CV

UVU:通过视觉-语言统一自回归范式提升多模态理解

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大语言模型细粒度视觉理解受限于稀疏文本监督的问题,本文提出UVU框架,在预训练阶段引入连续视觉编码与像素级码本,实现像素与文本标记统一自回归生成,显著提升多模态理解性能。

中文摘要 AI 辅助

尽管多模态大语言模型(MLLMs)取得了显著进展,但其细粒度视觉理解仍受限于主要依赖稀疏的文本监督。现有的引入视觉监督的努力通常在后训练阶段进行,此时视觉表示已基本固定,导致此类信号主要作为辅助约束,而非塑造感知特征的主要力量。在本文中,我们旨在通过将视觉监督直接纳入预训练阶段,从根本上重塑模型的感知骨干。我们观察到,像素级图像块和文本标记天然共存于一个共享的、原始的高维空间中,该空间具有固有的输入对称性。基于这一洞察,我们提出了UVU,一种新颖的视觉-语言统一自回归框架,该框架摒弃了向量量化。它独特地采用连续视觉编码以实现视觉输入的无损表示,并提出一种大规模迭代层次聚类算法来构建像素级视觉码本,从而扩展统一监督的词汇表,使像素级图像标记与文本标记能够自回归生成。UVU有效协同像素级视觉感知与语义级视觉理解,内化视觉重建能力,并释放视觉监督在预训练阶段增强理解的促进作用。跨多个任务的广泛实验表明,在UVU的监督学习范式下,MLLMs能够实现更优的多模态理解性能。

英文摘要

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations have already been largely fixed, causing such signals to act mainly as auxiliary constraints rather than as a primary force for shaping perceptual features. In this paper, we aim to fundamentally reshape the model's perceptual backbone by incorporating vision supervision directly into the pre-training stage. We observe that pixel-level image patches and textual tokens naturally coexist in a shared, raw high-dimensional space characterized by an inherent input symmetry. Leveraging this insight, we propose UVU, a novel vision-language unified autoregressive framework that eschews vector quantization. It uniquely employs continuous visual encoding for lossless representation of visual inputs and proposes a large-scale iterative hierarchical clustering algorithm to construct a pixel-level visual codebook, thereby extending the vocabulary for unified supervision and enabling autoregressive generation of pixel-level image tokens alongside textual tokens. UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage. Extensive experiments across multiple tasks demonstrate that MLLMs are capable of achieving superior multimodal understanding performance under the supervised learning paradigm of UVU.

发表机构

  • Tsinghua University(清华大学)
  • Tencent Youtu Lab(腾讯优图实验室)
  • Nanjing University(南京大学)
  • University of Glasgow(格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑