RA-CFGCache:从分支级准则到无分类器引导下的引导风险控制
RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance
浏览论文内容
中文总结 AI 辅助
针对扩散模型无分类器引导下缓存重用误差控制问题,提出RA-CFGCache风险对齐缓存框架,通过CFG感知组合与传播感知重缩放,在多种模型上提升效率-保真度权衡。
中文摘要 AI 辅助
扩散模型能够生成高质量的视觉内容,但迭代去噪过程计算开销巨大,尤其是在需要同时进行条件评估和无条件评估的无分类器引导(CFG)下。免训练缓存通过重用先前计算的特征或预测来降低这一成本。然而,现有的分支局部重用准则并未明确考虑缓存错误在CFG下如何组合,以及局部扰动如何影响最终输出。我们识别出缓存控制中的两种错位:分支引导错位,即引导误差取决于分支误差的大小和方向对齐;以及局部-最终错位,即局部误差的下游影响在不同时间步上有所不同。我们提出RA-CFGCache,一种在CFG下结合这两个因素的风险对齐缓存框架,同时保持采样调度和引导规则不变。CFG感知的引导风险组合利用CFG系数和离线校准的跨分支对齐来组合现有的分支级代理。传播感知重缩放进一步用从隔离重用扰动中校准的时间步相关传播先验对所得引导风险估计进行加权。在线阈值控制器随后决定何时联合刷新或重用两个分支。在FLUX.1-dev、Wan2.1-T2V-1.3B和CogVideoX-2B上的实验表明,与评估的免训练缓存基线相比,效率-保真度权衡得到改善。此外,RA-CFGCache兼容多种基础代理族,包括TeaCache-、DiCache-和MagCache风格的估计器,并在几乎不变的延迟下持续提高保真度。代码可在该https URL获取。
英文摘要
Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency--fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at https://github.com/yiming-l21/RA-CFGCache.git.
发表机构
- Tsinghua University(清华大学)
- BNRist(北京国家研究中心)
机构由 AI 辅助整理,请以论文原文为准。