arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReaDiT Guidance:基于扩散Transformer特征的图像与视频生成控制方法

ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

Jay Mahajan, Chang Liu, Rauf Makharov, Viraj Shah, Alexander Schwing, Svetlana Lazebnik

arXiv 2609.04649首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ReaDiT Guidance轻量框架,利用DiT块特征控制图像/视频生成,可实现空间目标及相机运动控制,参数更少且效果优于同类方法。

AI 中文摘要

我们提出了DiT读出(ReaDiT) Guidance,这是一种轻量级框架,用于通过扩散Transformer(DiT)模型的内部特征表示来控制生成过程。ReaDiT Guidance利用单个DiT块的特征,根据测试时提供的空间目标(如深度、姿态或边缘图)引导生成过程。此外,由于现代文本到视频模型大多基于DiT骨干构建,ReaDiT Guidance可自然扩展到视频生成,实现相机与运动控制。实验结果表明,与现有基于特征的方法及现成适配器方法相比,我们的方法在参数更少的情况下达到了相当或更优的结果。

英文摘要

We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Transformer (DiT) models via their internal feature representations. ReaDiT Guidance uses features from a single DiT block to steer the generative process according to spatial targets - like depth, pose, or edge maps - provided at test time. Furthermore, since modern text-to-video models are largely built on DiT backbones, ReaDiT Guidance naturally extends to video generation, enabling camera and motion control. Experimental results demonstrate that our approach achieves competitive or improved results compared to existing feature-based and off-the-shelf adapter-based approaches while requiring fewer parameters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑