发表机构
The University of Hong Kong; Tsinghua University; Harbin Engineering University; University of Newcastle; Shenzhen University(香港大学; 清华大学; 哈尔滨工程大学; 纽卡斯尔大学; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出频谱张力诊断与频谱传输稳态调节方法,用于识别并校正视频生成中时间状态传输的失衡,提升时间一致性与视觉质量。
AI 中文摘要
可靠的视频生成不仅需要高质量的画面来构成连贯的故事,模型还必须维持一个持久的状态,将身份、场景布局、运动和细节等视觉属性随时间进行传输。现有的免训练方法主要强化跨帧注意力或分析局部注意力熵,但这些视角无法揭示时间交互是否处于健康的传输状态。在本工作中,我们通过时间状态传输的视角研究视频生成。我们引入了频谱张力,一种有符号的诊断指标,将局部注意力扩散性与全局频谱多样性进行比较,并利用它识别两种相反的时间故障:碎片化传输和过度混合热点。基于这一诊断,我们提出了频谱传输稳态,一种免训练的调节器,可温和地校正病态时间状态,同时基本保留平衡状态。在预训练视频生成模型上的实验表明,原始模型常处于失衡的时间状态,而我们的方法选择性地对最差的时间热点施加更大的校正,并在无需微调的情况下提升时间一致性和视觉质量。代码:此 https URL
英文摘要
Reliable video generation requires more than high-quality frames to form a coherent story: a model must maintain a persistent state, transporting visual attributes such as identity, scene layout, motion, and fine details across time. Existing training-free methods mainly strengthen cross-frame attention or analyze local attention entropy, but these views do not reveal whether temporal interactions stay in a healthy transport regime. In this work, we study video generation through the perspective of Temporal State Transport. We introduce Spectral Tension, a signed diagnostic that compares local attention diffuseness with global spectral diversity, and use it to identify two opposite temporal failures: fragmented transport and over-mixing hotspots. Based on this diagnosis, we propose Spectral Transport Homeostasis, a training-free regulator that softly corrects pathological temporal states while largely preserving balanced ones. Experiments on pretrained video generation models show that the original model often occupies imbalanced temporal regimes, whereas our method selectively applies larger corrections to the worst temporal hotspots and improves temporal consistency and visual quality without finetuning. Code: https://github.com/lytang63/temporal-state-transport
Comments**Best Paper** Award! ICML 2026 F2S Workshop