通过注意力头控制扩散变换器中的运动传递
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
- Yonsei University(延世大学)
- LG Electronics(LG电子)
- University of California, Merced(加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究如何控制扩散变换器中的运动传递,通过在注意力头层面分析,提出无需参数更新的头感知可控运动传递框架,利用语义对应引导和选择性特征注入,实现精确运动传递并为可控视频生成提供可解释基础。
AI中文摘要:
扩散变换器(DiTs)在高质量、时间连贯的视频生成方面取得了进展。然而,将其扩展到运动传递仍然具有挑战性,因为对DiTs中的运动和结构表示理解有限。我们在注意力头层面分析视频DiTs,识别出专门用于运动和空间结构的不同头。基于此,我们提出了一个无需参数更新的头感知可控运动传递框架。该方法通过语义对应引导从运动专用头中提炼运动线索,并通过选择性特征注入保留结构。这种头级控制不仅实现了精确的运动传递,还为使用DiTs进行可控视频生成提供了可解释的基础。
英文摘要:
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.