UniCode:学习多模态大语言模型的统一码本
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
- Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
- School of Computer Science Peking University(北京大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
UniCode通过统一码本高效标记化视觉与文本信号,结合语言驱动迭代训练和图像解压缩预训练,实现高质量图像生成,并在VQA基准上达到领先性能。
AI中文摘要:
在本文中,我们提出了UniCode,一种在多模态大语言模型(MLLMs)领域中的新颖方法,它学习一个统一的码本来高效地标记化视觉、文本以及其他类型的信号。这项创新解决了现有MLLMs的一个关键局限性:它们依赖于仅文本的码本,这限制了MLLM在多模态上下文中生成图像和文本的能力。为此,我们提出了一种语言驱动的迭代训练范式,并配合一个我们称之为“图像解压缩”的上下文内预训练任务,使我们的模型能够解释压缩的视觉数据并生成高质量的图像。统一的码本使我们的模型能够将视觉指令调整扩展到非语言生成任务。此外,UniCode可适应多种堆叠量化方法,以将视觉信号压缩成更紧凑的标记表示。尽管在训练期间使用的参数和数据显著更少,UniCode在视觉重建和生成方面展示了有前景的能力。它还在各种VQA基准测试中取得了与领先MLLMs相当的性能。
英文摘要:
In this paper, we propose \textbf{UniCode}, a novel approach within the domain of multimodal large language models (MLLMs) that learns a unified codebook to efficiently tokenize visual, text, and potentially other types of signals. This innovation addresses a critical limitation in existing MLLMs: their reliance on a text-only codebook, which restricts MLLM's ability to generate images and texts in a multimodal context. Towards this end, we propose a language-driven iterative training paradigm, coupled with an in-context pre-training task we term ``image decompression'', enabling our model to interpret compressed visual data and generate high-quality images.The unified codebook empowers our model to extend visual instruction tuning to non-linguistic generation tasks. Moreover, UniCode is adaptable to diverse stacked quantization approaches in order to compress visual signals into a more compact token representation. Despite using significantly fewer parameters and less data during training, Unicode demonstrates promising capabilities in visual reconstruction and generation. It also achieves performances comparable to leading MLLMs across a spectrum of VQA benchmarks.