arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TSGate:面向扩散 Transformer 的时间步感知门控注意力

TSGate: Timestep-Aware Gated Attention for Diffusion Transformers

Boyu Zhang, Yifan Liu, Shuxia Lin, Qingjian Ni, Yinfei Xu, Xu Yang

arXiv 2609.34539首次发表:更新:

发表机构

Alibaba Token Hub, Alibaba Group; Southeast University(阿里巴巴集团,阿里巴巴Token Hub; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散 Transformer 在域外提示下生成质量下降的问题,本文提出时间步感知门控注意力(TSGate),通过注入时间步条件偏置自适应调整门控行为,在多个基准上显著提升生成质量,原始提示 DPG 分数提高 9.5%。

AI 中文摘要

扩散 Transformer(DiTs)已成为高保真图像和视频生成的主流架构。最近的 DiT 系统越来越多地使用结构化提示进行训练,从而提高了标题质量和提示遵循能力。然而,在域外(OOD)提示(包括用户在推理时提供的自由形式描述)下,其生成质量可能会严重下降。尽管基于 LLM 的重写可以将这些提示转换为结构化格式,但它并不能保证重写后的提示与训练分布一致。我们的分析将这种质量下降归因于注意力汇聚(attention sinks)以及早期步骤中图像到文本注意力的减少,并表明仅抑制注意力汇聚不足以恢复生成质量。尽管有效抑制了注意力汇聚,但使用标准门控注意力训练的模型在早期步骤中仍表现出图像到文本注意力的减少和次优的生成质量。基于这些见解,我们提出了时间步感知门控注意力(TSGate),它将一个时间步条件偏置注入门控信号中,从而使门控行为在去噪步骤中自适应调整。大量实验表明,TSGate 在多个基准上始终优于基线和标准门控注意力,将原始提示的 DPG 分数相对基线提高了 9.5%。

英文摘要

Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prompts, including the free-form descriptions supplied by users at inference time. Although LLM-based rewriting can convert these prompts into structured formats, it does not guarantee that the rewritten prompts align with the training distribution. Our analysis links this degradation to attention sinks and reduced early-step image-to-text attention and shows that sink suppression alone is insufficient to restore generation quality. Despite effective sink suppression, models trained with standard gated attention exhibit reduced early-step image-to-text attention and suboptimal generation quality. Based on these insights, we propose Timestep-Aware Gated Attention (TSGate), which injects a timestep-conditioned bias into the gate signal so that gating behavior adapts across denoising steps. Extensive experiments show that TSGate consistently outperforms both the baseline and standard gated attention across multiple benchmarks, improving the raw-prompt DPG score by 9.5% over the baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑