arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07815cs.CVcs.AIcs.CL

VoT:视觉思维用于统一多模态表示对齐

VoT: Vision-of-Thought for Unified Multimodal Representation Alignment

  • ByteDance Seed(字节跳动Seed)

机构由 AI 辅助整理,请以论文原文为准。

Jingxiang Sun, Chao Liao, Zhengxiong Luo, Chaorui Deng, Chen-lin Zhang, Junke Wang, Ceyuan Yang, Haoqi Fan, Weilin Huang

AI总结:

本文提出VoT框架,在视觉语言模型和扩散变换器间引入离散视觉思维层,通过闭环目标训练分词器,生成高层视觉计划标记,以改善文本到图像生成的语义对齐和可控性。

AI中文摘要:

当前的文本到图像系统通常采用“文本编码器加扩散解码器”的范式,其中文本语义直接调制连续潜在噪声。尽管这些方法取得了成功,但它们缺乏一种显式的、可解释的中间表示,以有效桥接高层语言语义和低层视觉信号。在本文中,我们提出了视觉思维(VoT),一个在视觉语言模型(VLMs)和扩散变换器(DiTs)之间引入离散视觉思维层的框架。我们不将VLM仅视为文本编码器,而是将其用作多模态规划器,在渲染像素之前生成表示高层视觉计划(如对象和布局)的离散VoT标记。我们在VLM语义空间中训练一个专门的VoT分词器,采用结合VLM对齐、特征重建和向量量化损失的闭环目标。这些目标使标记在语义上可被VLM读取,同时保留生成所需的视觉信息。实验结果表明,VoT改善了语义对齐,并为可解释和可控生成提供了结构化接口。

英文摘要:

Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their success, these methods lack an explicit, interpretable intermediate representation that effectively bridges high-level linguistic semantics and low-level visual signals. In this paper, we propose Vision-of-Thought (VoT), a framework that introduces a discrete visual-thinking layer between vision-language models (VLMs) and diffusion transformers (DiTs). Instead of treating VLMs merely as text encoders, we use them as multimodal planners that generate discrete VoT tokens representing high-level visual plans, such as objects and layouts, before rendering pixels. We train a specialized VoT tokenizer in the VLM semantic space with a closed-loop objective that combines VLM alignment, feature reconstruction, and vector-quantization losses. These objectives make the tokens semantically readable by the VLM while preserving the visual information needed for generation. Experimental results demonstrate that VoT improves semantic alignment and provides a structured interface for interpretable and controllable generation.

↑