SelfLift:通过自恢复分辨率转换加速少步扩散模型
SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出SelfLift框架,通过自恢复分辨率转换减少少步扩散模型的端到端延迟,在FLUX.2-Klein等数据集上实现41.5%至44.1%的延迟降低,结合时间步蒸馏后加速比达19.21倍至29.61倍。
AI中文摘要:
少步扩散模型大幅压缩了时序计算量,使得每次模型评估的空间成本成为推理延迟的主导来源。渐进分辨率推理通过在低分辨率下执行早期去噪、将高分辨率计算留作精细化处理来降低该成本。然而,现有方法通常直接提升中间隐变量,并依赖后续步骤吸收由此产生的分布不匹配问题。在少步场景下,有限的恢复预算会使这些误差成为可见的伪影,限制了转换发生的时机,进而制约了其执行效率。我们提出SelfLift,一种自恢复渐进分辨率框架,可从生成模型本身推导转换修复信号和轨迹对齐的监督信号。SelfLift-zero提出了一种无需训练的伪影感知一致性提升方法,利用直接隐变量提升与像素VAE重编码之间的分歧,作为局部伪影风险信号和模型原生修正方向,无需外部超分辨率、额外去噪器评估或采样 schedule 修改即可实现可靠的后期转换。在此稳健转换基础上,SelfLift-rich在学生访问的状态上执行策略内自恢复,从内部自教师转移密集高分辨率指导,同时与改变后的渐进分辨率动态保持一致。在FLUX.2-Klein和Z-Image-Turbo上,SelfLift分别将端到端延迟降低41.5%和44.1%,结合时间步长蒸馏后,与对应的50步模型相比,分别实现29.61倍和19.21倍的整体加速,同时保持了竞争力的生成质量,为少步扩散模型建立了更强的速度-质量前沿。
英文摘要:
Few-step diffusion models substantially compress temporal computation, making the spatial cost of each model evaluation an increasingly dominant source of inference latency. Progressive-resolution inference reduces this cost by performing early denoising at low resolution and reserving high-resolution computation for refinement. However, existing methods typically lift intermediate latents directly and rely on subsequent steps to absorb the induced distribution mismatch. In the few-step regime, the limited recovery budget leaves these errors as visible artifacts, constraining how late the transition can occur and, consequently, how efficiently it can be performed. We introduce SelfLift, a self-recovering progressive-resolution framework that derives both transition-repair signals and trajectory-aligned supervision from the generative model itself. SelfLift-zero proposes a training-free Artifact-Aware Consistency Lift, using disagreement between direct latent lifting and pixel-VAE re-encoding as both a localized artifact-risk signal and a model-native correction direction. It enables reliable late transitions without external super-resolution, extra denoiser evaluations, or sampling-schedule modifications. Building on this robust transition, SelfLift-rich performs On-Policy Self Recovery on student-visited states, transferring dense high-resolution guidance from an internal self-teacher while remaining aligned with the altered progressive-resolution dynamics. Across FLUX.2-Klein and Z-Image-Turbo, SelfLift reduces end-to-end latency by 41.5% and 44.1%, respectively. Combined with timestep distillation, it delivers overall speedups of 29.61x and 19.21x over the corresponding 50-step models while preserving competitive generation quality, establishing a stronger speed-quality frontier for few-step diffusion.