发表机构
Westlake AGI Lab(西湖AGI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BAG(预算感知门控机制),通过离线到在线调度蒸馏训练轻量门控网络,在FLUX.1-dev和Wan2.1上实现优于现有最优缓存方法的性能。
AI 中文摘要
扩散缓存是一种通过在去噪步骤间复用中间特征来加速扩散Transformer(DiTs)的轻量策略,但现有范式存在根本权衡:在线启发式方法缺乏全局预算感知,而静态调度方案缺乏实例适应性,无法灵活适配不同的运行时预算约束。为填补这一空白,我们提出BAG(Budget-Aware Gating,预算感知门控机制),一种新型缓存策略,将全局预算 pacing 与动态、实例自适应的特征复用相统一。BAG不依赖人工规则,而是采用轻量门控网络,通过联合以预算状态和局部轨迹反馈为条件,在每一步动态决定执行完整计算还是复用缓存特征。我们通过离线到在线的调度蒸馏训练该策略,将离线搜索得到的调度决策迁移至紧凑的在线门控网络。在FLUX.1-dev和Wan2.1上开展的大量实验表明,BAG在各类加速层级上均始终优于当前最优的缓存方法,且在不同分辨率、随机种子和引导尺度下保持鲁棒性。代码将公开。
英文摘要
Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but existing paradigms face a fundamental trade-off: online heuristics lack global budget awareness, whereas static schedules lack instance adaptivity and fail to flexibly adapt to varying runtime budget constraints. To bridge this gap, we present BAG (Budget-Aware Gating), a novel caching policy that unifies global budget pacing with dynamic, instance-adaptive feature reuse. Rather than relying on hand-crafted rules, BAG employs a lightweight gating network that dynamically decides whether to execute a full computation or reuse cached features at each step by jointly conditioning on the budget state and local trajectory feedback. We train this policy via offline-to-online schedule distillation, transferring the decision-making of offline-searched schedules into a compact online gate. Extensive experiments on FLUX.1-dev, Wan2.1, and Qwen-Image-2512 demonstrate that BAG consistently outperforms state-of-the-art caching methods across various speedup tiers while remaining robust across different resolutions, seeds, and guidance scales. Code will be released.
Comments23 pages, 13 figures, and 14 tables. Code Link: see AGI-Lab/BAG" target="_blank" rel="noopener">https://github.com/Westlake-AGI-Lab/BAG