发表机构
Shanghai Jiao Tong University; Shandong University; Terminal Intelligent Computing Division, Alibaba Cloud; Jilin University; Xi’an Jiaotong University(上海交通大学; 山东大学; 阿里云终端智能计算事业部; 吉林大学; 西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对扩散变换器动态分辨率采样的冗余计算问题,提出AViTS框架,通过时空重要性感知的选择性上采样减少冗余,在FLUX等模型上实现显著加速且与其他优化技术兼容
AI 中文摘要
扩散变换器(DiTs)可实现高质量生成,但因迭代采样而计算成本高昂。动态分辨率采样通过在低分辨率下去噪降低早期阶段的成本;然而,在分辨率转换时对所有潜在令牌进行统一上采样会导致冗余计算,且可能降低细节一致性。现有的部分上采样策略通常依赖局部潜在结构线索或单步统计,难以联合捕捉令牌-文本语义相关性以及扩散步骤间的令牌级表示动态。我们提出AViTS,一种用于动态分辨率DiTs的自适应时空令牌选择框架。AViTS通过潜在-文本注意力建模空间重要性,通过扩散时间步间的令牌级特征变化建模时间重要性,并将二者融合以实现感知时空重要性的选择性上采样:它优先对关键令牌进行分辨率优化,同时推迟处理较不重要的令牌,从而减少冗余的高分辨率计算并改善质量-效率权衡。AViTS在FLUX上实现了最高6.34倍的推理加速,在Qwen-Image-Edit和FLUX.1-Kontext-dev上实现了近9倍的FLOPs减少,且与蒸馏、量化和特征缓存技术正交,结合蒸馏模型可达到14.76倍的加速。代码:this https URL
英文摘要
Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low resolution; however, uniformly upsampling all latent tokens at resolution transitions incurs redundant computation and may degrade fine-detail consistency. Existing partial upsampling strategies typically rely on local latent structure cues or single-step statistics, making it difficult to jointly capture token-text semantic relevance and token-wise representation dynamics across diffusion steps. We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic-resolution DiTs. AViTS models spatial importance via latent-text attention and temporal importance via token-level feature variation across diffusion timesteps, and fuses them to enable spatiotemporal importance-aware selective upsampling: it prioritizes resolution refinement for critical tokens while deferring less important ones, thereby reducing redundant high-resolution computation and improving the quality-efficiency trade-off. AViTS achieves up to 6.34x on FLUX and nearly 9x FLOPs reduction on Qwen-Image-Edit and FLUX.1-Kontext-dev, orthogonal to distillation, quantization, and feature caching, and reaching 14.76x with distilled models. Code: https://github.com/QHR69/AViTS
CommentsAccepted to ECCV 2026. 20 pages including appendix. Code: https://github.com/QHR69/AViTS