arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

研究统一多模态模型中的图像分词器作为视觉语言

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang Liu

arXiv 2609.09143首次发表:更新:

发表机构

Amazon FAR; University of Washington(亚马逊FAR; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建纯自回归测试平台,通过任务特定损失分析图像分词器在联合多模态训练中的行为,发现损失需按任务分析且I2T损失更一致,并揭示重建质量与性能非正相关及分词器选择影响文本建模。

AI 中文摘要

图像分词器定义了统一多模态模型的“视觉语言”,然而通常通过孤立的指标或仅生成/仅理解的评估来研究。这些评估未能完全捕捉视觉标记在与文本联合建模时的行为。我们构建了一个受控的纯自回归测试平台,并在多模态持续预训练过程中,跨文本、图像、文本到图像(T2I)和图像到文本(I2T)预测跟踪任务特定的验证损失。我们研究这些损失如何缩放以及它们与下游性能的关系,然后利用它们来研究多模态可学习性——即图像和文本标记如何被联合建模——以及分词器设计。我们发现:(1)损失应按任务进行分析,因为它们表现出不同的缩放行为并对分词器进行不同排序。(2)损失与性能的关系取决于预测的标记空间:对于固定的分词器,T2I和I2T损失与生成质量相关,但在不同分词器之间,T2I损失与性能的关系随图像标记空间变化,而I2T损失在共享文本词汇表上计算,提供了更一致的信号。I2T损失在监督微调后也与生成和视觉理解性能相关。利用损失作为视角,我们表明(3)更好的重建不一定产生更低的任务特定损失或更强的下游性能,并且(4)图像分词器的选择在联合优化下可能影响文本建模。作为案例研究,我们重新审视三个分词器设计轴——判别器、语义监督和词汇表大小——以检查它们对联合建模和下游性能的影响。总之,我们的测试平台为图像分词器作为视觉语言提供了互补视角,强调了它们在联合多模态训练中与文本的相互作用。

英文摘要

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.

Comments27 pages, 23 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑