UVU:通过视觉-语言统一自回归范式提升多模态理解
UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm
浏览论文内容
中文总结 AI 辅助
针对多模态大语言模型细粒度视觉理解受限于稀疏文本监督的问题,本文提出UVU框架,在预训练阶段引入连续视觉编码与像素级码本,实现像素与文本标记统一自回归生成,显著提升多模态理解性能。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLMs)取得了显著进展,但其细粒度视觉理解仍受限于主要依赖稀疏的文本监督。现有的引入视觉监督的努力通常在后训练阶段进行,此时视觉表示已基本固定,导致此类信号主要作为辅助约束,而非塑造感知特征的主要力量。在本文中,我们旨在通过将视觉监督直接纳入预训练阶段,从根本上重塑模型的感知骨干。我们观察到,像素级图像块和文本标记天然共存于一个共享的、原始的高维空间中,该空间具有固有的输入对称性。基于这一洞察,我们提出了UVU,一种新颖的视觉-语言统一自回归框架,该框架摒弃了向量量化。它独特地采用连续视觉编码以实现视觉输入的无损表示,并提出一种大规模迭代层次聚类算法来构建像素级视觉码本,从而扩展统一监督的词汇表,使像素级图像标记与文本标记能够自回归生成。UVU有效协同像素级视觉感知与语义级视觉理解,内化视觉重建能力,并释放视觉监督在预训练阶段增强理解的促进作用。跨多个任务的广泛实验表明,在UVU的监督学习范式下,MLLMs能够实现更优的多模态理解性能。
英文摘要
Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations have already been largely fixed, causing such signals to act mainly as auxiliary constraints rather than as a primary force for shaping perceptual features. In this paper, we aim to fundamentally reshape the model's perceptual backbone by incorporating vision supervision directly into the pre-training stage. We observe that pixel-level image patches and textual tokens naturally coexist in a shared, raw high-dimensional space characterized by an inherent input symmetry. Leveraging this insight, we propose UVU, a novel vision-language unified autoregressive framework that eschews vector quantization. It uniquely employs continuous visual encoding for lossless representation of visual inputs and proposes a large-scale iterative hierarchical clustering algorithm to construct a pixel-level visual codebook, thereby extending the vocabulary for unified supervision and enabling autoregressive generation of pixel-level image tokens alongside textual tokens. UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage. Extensive experiments across multiple tasks demonstrate that MLLMs are capable of achieving superior multimodal understanding performance under the supervised learning paradigm of UVU.
发表机构
- Tsinghua University(清华大学)
- Tencent Youtu Lab(腾讯优图实验室)
- Nanjing University(南京大学)
- University of Glasgow(格拉斯哥大学)
机构由 AI 辅助整理,请以论文原文为准。