arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31654cs.CVcs.AI

视频扩散训练过程中的时间注意力头特化

Temporal-Attention Head Specialization During Video Diffusion Training

Taewoo Ha, Shafayat Mowla Anik, Dae Yeol Lee, Byeong Kil Lee, Jeeho Ryoo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过检查点解析的逐头普查,发现视频扩散训练中仅少数时间注意力头(4-13%)特化,且集中于首个时间块,并收敛到局部帧路由模式,但未发现与视频质量的因果联系,同时提供了一种可迁移的分析方法。

中文摘要 AI 辅助

视频扩散Transformer依赖时间注意力来协调跨帧信息,然而关于这一机制几乎所有的已知内容都来自对已训练模型的分析,因此训练过程中时间注意力结构何时形成以及在哪里形成仍然缺乏表征。群体平均值也可能掩盖这一现象,因为少数特化的注意力头与多数扩散的注意力头在均值上相互抵消。为此,我们对跨越三个模型规模(306M至1.03B参数)的九次Open-Sora STDiT训练运行中的每一个时间注意力头进行了检查点解析的全量普查,并使用熵归一化的跨帧注意力集中度(CFAC)度量对每个头进行评分,同时采用预注册的变点与效应量选择规则。该普查揭示了平均值所掩盖的稀疏图景。在每次运行中,聚合CFAC持平或下降,而少数注意力头(在全网格运行中约占4%至13%)发展出显著的集中性。跨种子(随机种子)而言,可复现的信号是位置性的但属于块级别。选中的头反复出现在第一个时间块中,而在考虑块成员关系后,单个头的坐标并不复现。在分析的760M选中的头中,注意力图收敛到一小类局部帧路由模式,即自帧对角线和相邻帧带状模式,即使负责的坐标在不同运行中有所不同。相关性和消融分析并未建立与生成视频质量的因果联系,因此我们相应地对我们的主张加以限定。除了这一STDiT系列之外,本研究还贡献了一种可迁移的方法论。在固定选择规则下进行检查点解析的逐头分析,可以揭示其他分解式视频扩散Transformer中的稀疏时间组织,并且通过适配的路由度量,也可用于联合时空架构。

英文摘要

Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remains poorly characterized. Population averages can also hide it, since a few specializing heads and a diffusing majority cancel in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters), scoring each head with an entropy-normalized measure of cross-frame attention concentration (CFAC) under a preregistered change-point and effect-size selection rule. The census reveals the sparse picture that averages obscure. Aggregate CFAC is flat or decreasing in every run, while a small minority of heads, roughly 4--13% in full-grid runs, develops pronounced concentration. Across seeds, the reproducible signal is positional but block-level. Selected heads repeatedly arise in the first temporal block, whereas individual head coordinates do not reproduce once block membership is accounted for. Among the analyzed 760M selected heads, attention maps converge to a small repertoire of local frame-routing motifs, self-frame diagonals and adjacent-frame bands, even when the responsible coordinates differ across runs. Correlation and ablation analyses do not establish a causal link to generated video quality, and we bound our claims accordingly. Beyond this STDiT family, the study contributes a transferable methodology. Checkpoint-resolved, per-head analysis under fixed selection rules can expose sparse temporal organization in other factorized video diffusion transformers and, with adapted routing metrics, in joint spatio-temporal architectures.

发表机构

  • University of Colorado - Colorado Springs(科罗拉多大学科罗拉多斯普林斯分校)
  • Dolby Laboratories(杜比实验室)
  • Fairleigh Dickinson University(菲尔莱狄更斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑