arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19139cs.CV

文本模板令牌是扩散Transformer中的隐式语义寄存器

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

  • Nanjing University(南京大学)
  • Alibaba Group(阿里巴巴集团)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang

AI总结:

研究文本到图像扩散Transformer中内部计算,引入因果可解释性框架,发现结构模板令牌充当隐式语义寄存器,据此设计剪枝规则,还揭示其生成计算组织方式,提供了DiTs内部机制的因果视图。

AI中文摘要:

文本到图像的扩散Transformer(DiTs)联合处理文本和图像令牌,但其去噪过程中的内部计算仍知之甚少。我们为现代大规模DiTs引入了一个因果可解释性框架,将注意力分解与跨令牌跨度、头和层的定向干预相结合。通过将提示内容令牌与结构模板令牌分离,发现结构令牌在编码器输出时携带的特定提示信息很少。令人惊讶的是,它们成为主要的图像到文本注意力汇聚点,并在DiT中因果性地维持对象身份,充当隐式语义寄存器。它们通过提示语义先注入图像潜空间然后再读回模板令牌来间接获得这种身份。受此启发,我们为DiTs设计了一个无需训练的剪枝规则。对提示令牌关注最强的头是可舍弃的,剪枝它们可去除20%的注意力FLOP,在GenEval上仅下降1.4分。我们还揭示了DiTs中的生成计算是如何跨头和深度组织的,将语义路由与视觉合成分离,并从身份形成发展到传播和细化。我们的工作不仅揭示了在输入时编码语义的令牌不一定是在生成过程中维持语义的令牌,还提供了DiTs内部机制的因果视图。

英文摘要:

Modern text-to-image diffusion transformers (DiTs) generate images through joint attention, in which text and image tokens interact directly within a single sequence. In large-scale DiTs, the conditioning input contains not only the user prompt but also chat-template tokens introduced by LLM-based text encoders. Yet how these tokens participate in the denoising computation remains poorly understood. To probe this, we introduce a causal interpretability framework. Using it to separate prompt-content tokens from chat-template tokens, we find that the template tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly. Rather than reading the prompt tokens, they draw the identity from the image latents into which the prompt semantics have already been injected at the very first layer. We further reveal a division of labor across heads and depth in DiTs, where distinct heads route semantics or render visual structure, and identity is committed in early blocks, carried by middle blocks, and refined in late ones. As a practical payoff, this analysis yields a training-free pruning rule that removes the causally inert prompt-reading heads and cuts $20\%$ of joint-attention FLOPs at a $1.4$-point cost in GenEval accuracy. Overall, our work not only reveals that the tokens encoding semantics at the input need not be those that maintain them during generation, but also provides a causal view of internal mechanisms in diffusion transformers.

↑