VoT:视觉思维用于统一多模态表示对齐
VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
- ByteDance Seed(字节跳动Seed)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出VoT框架,在视觉语言模型和扩散变换器间引入离散视觉思维层,通过闭环目标训练分词器,生成高层视觉计划标记,以改善文本到图像生成的语义对齐和可控性。
AI中文摘要:
当前的文本到图像系统通常采用“文本编码器加扩散解码器”的范式,其中文本语义直接调制连续潜在噪声。尽管这些方法取得了成功,但它们缺乏一种显式的、可解释的中间表示,以有效桥接高层语言语义和低层视觉信号。在本文中,我们提出了视觉思维(VoT),一个在视觉语言模型(VLMs)和扩散变换器(DiTs)之间引入离散视觉思维层的框架。我们不将VLM仅视为文本编码器,而是将其用作多模态规划器,在渲染像素之前生成表示高层视觉计划(如对象和布局)的离散VoT标记。我们在VLM语义空间中训练一个专门的VoT分词器,采用结合VLM对齐、特征重建和向量量化损失的闭环目标。这些目标使标记在语义上可被VLM读取,同时保留生成所需的视觉信息。实验结果表明,VoT改善了语义对齐,并为可解释和可控生成提供了结构化接口。
英文摘要:
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their success, these methods lack an explicit, interpretable intermediate representation that effectively bridges high-level linguistic semantics and low-level visual signals. In this paper, we propose Vision-of-Thought (VoT), a framework that introduces a discrete visual-thinking layer between vision-language models (VLMs) and diffusion transformers (DiTs). Instead of treating VLMs merely as text encoders, we use them as multimodal planners that generate discrete VoT tokens representing high-level visual plans, such as objects and layouts, before rendering pixels. We train a specialized VoT tokenizer in the VLM semantic space with a closed-loop objective that combines VLM alignment, feature reconstruction, and vector-quantization losses. These objectives make the tokens semantically readable by the VLM while preserving the visual information needed for generation. Experimental results demonstrate that VoT improves semantic alignment and provides a structured interface for interpretable and controllable generation.