学习读取扩散Transformer中的上下文令牌
Learning to Read the Contextual Tokens in Diffusion Transformers
浏览论文内容
中文总结 AI 辅助
本文提出通过轻量级瓶颈网络将扩散Transformer的上下文令牌映射到冻结LLM进行自然语言解读,发现其编码全局场景语义,并据此提出上下文对齐训练技术以提升生成质量。
中文摘要 AI 辅助
多模态扩散Transformer(MM-DiTs)在生成过程中联合处理视觉和文本表示。这些模型通过多模态注意力反复更新文本令牌,形成动态的上下文令牌,其功能尚不明确。在这项工作中,我们引入了一个通过自然语言质询来读取这一上下文空间的框架。我们训练了一个轻量级瓶颈网络,将中间上下文令牌映射到冻结的大型语言模型(LLM)的输入空间中,使LLM能够直接基于这些隐藏表示回答关于生成图像的问题。我们的读取器揭示,上下文令牌编码了生成场景的丰富全局表示:包括提示中未明确指定的属性在内的生成特定语义,在去噪过程中出人意料地早期即可访问,而随着时间推移,越来越细粒度的细节变得可读。值得注意的是,即使MM-DiT接收到空提示,这些信息仍然可解码,表明上下文令牌从不断演化的视觉表示本身积累了大量的图像特定信息。我们进一步发现,具有更可读上下文表示的生成结果往往获得更高的人类偏好评分。基于这些观察,我们引入了上下文对齐(Contextual Alignment),这是一种训练技术,显式强化上下文令牌中编码的视觉-语义信息,从而提高生成质量和分布覆盖。综合来看,我们的结果确立了上下文令牌既可作为MM-DiTs内部动态的可解释视图,也可作为改进生成模型的有效目标。
英文摘要
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.
发表机构
- Tel Aviv University(特拉维夫大学)
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。