RefAdapt-DiT:用于参考条件扩散Transformer的自适应联合注意力
RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers
浏览论文内容
中文总结 AI 辅助
提出RefAdapt,一种免训练框架,通过结合目标准变化与目标对参考的注意力权重,在块粒度上自适应控制参考计算,实现最高3.54倍加速并保持视觉质量。
中文摘要 AI 辅助
扩散Transformer(DiTs)已成为高质量生成建模的标准骨干架构,然而在条件生成任务中部署它们仍计算开销巨大,因为双向联合注意力会反复处理大型参考流。现有优化方案虽能缓解通用的时间冗余,但通常依赖粗粒度的静态复用,忽视了参考与目标各自不同的动态特性。具体而言,我们观察到参考表示在生成轨迹中往往演化缓慢,而目标分配给参考的注意力权重通常很小;参考漂移与目标对参考的暴露程度共同决定了过时的参考状态对目标的影响程度。为利用这些模式,我们提出RefAdapt,一个免训练框架,用于自适应控制参考与目标之间的联合注意力。不同于僵化的静态策略,RefAdapt结合连续的目标准变化与先前观察到的目标对参考的注意力权重,在块粒度上自适应控制参考计算。在超少步数设置下,RefAdapt在4步MiniMax H3上实现高达2.097倍加速,在8步Qwen图像编辑上实现3.54倍加速,同时保持相当的视觉质量。
英文摘要
Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.