arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11081cs.CVcs.AI

通过注意力头控制扩散变换器中的运动传递

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

  • Yonsei University(延世大学)
  • LG Electronics(LG电子)
  • University of California, Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang

AI总结:

研究如何控制扩散变换器中的运动传递,通过在注意力头层面分析,提出无需参数更新的头感知可控运动传递框架,利用语义对应引导和选择性特征注入,实现精确运动传递并为可控视频生成提供可解释基础。

AI中文摘要:

扩散变换器(DiTs)在高质量、时间连贯的视频生成方面取得了进展。然而,将其扩展到运动传递仍然具有挑战性,因为对DiTs中的运动和结构表示理解有限。我们在注意力头层面分析视频DiTs,识别出专门用于运动和空间结构的不同头。基于此,我们提出了一个无需参数更新的头感知可控运动传递框架。该方法通过语义对应引导从运动专用头中提炼运动线索,并通过选择性特征注入保留结构。这种头级控制不仅实现了精确的运动传递,还为使用DiTs进行可控视频生成提供了可解释的基础。

英文摘要:

Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.

补充信息

↑