arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先脚手架后内化:扩散Transformer的表示注入

Scaffold Then Internalize: Representation Injection for Diffusion Transformers

Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen, Wei Liu, Qing Li, Xudong Mao

arXiv 2609.35292首次发表:更新:

发表机构

Sun Yat-sen University; Video Rebirth; The Hong Kong Polytechnic University(中山大学; Video Rebirth; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出REPI,通过将预训练编码器表示注入扩散Transformer并逐步内化,与REPA互补,显著加速训练,仅16万步即可匹配700万步的SiT性能。

AI 中文摘要

近期表示对齐(REPA)方法通过将扩散Transformer隐藏状态的投影与预训练视觉编码器的表示对齐,加速了扩散Transformer的训练。在本工作中,我们探索了与REPA相反且互补的方向:不是将扩散表示投影到编码器的空间中,而是将编码器表示注入到扩散Transformer中,使其积极参与去噪过程。为此,我们提出了表示注入(REPI),一种基于脚手架到内化策略的训练框架,其中投影的编码器表示最初作为临时脚手架,随后被扩散Transformer逐步内化。REPI在广泛的骨干网络上优于REPA,并且与REPA高度互补:将两者结合相比单独使用任一方法都能带来显著提升。值得注意的是,仅用16万训练步,REPI+REPA即可匹配训练了700万步的原始SiT,加速超过43.5倍。代码将在该https URL提供。

英文摘要

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over $43.5\times$. Code will be available at https://jeneveuxpas.github.io/REPI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑