arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IB-Flow:用于少步文本到图像生成的信息瓶颈引导的CFG蒸馏

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation

Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao

arXiv 2607.09133首次发表:更新:

发表机构

Tsinghua University; Wan Team, Alibaba Group; HKUST; Beijing Normal University(清华大学; 万团队,阿里巴巴集团; 香港科技大学; 北京师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对少步文本到图像生成中推理延迟问题,通过信息论将蒸馏建模为信息瓶颈引导的动态互信息博弈,提出双轨自适应框架,消除过度条件化伪影,在2步配置下实现SOTA生成保真度。

AI 中文摘要

虽然大规模文本到图像生成模型取得了前所未有的视觉性能,但对多步迭代求解器的依赖导致严重推理延迟。少步蒸馏成为流行的二维压缩范式,但现有框架受粗粒度盲目注入范式限制,无视图像生成的动态进化本质。我们通过信息论视角重新审视蒸馏过程,将其建模为受信息瓶颈原则约束的动态互信息博弈。提出双轨自适应框架,包括实例感知选择机制确定注入目标,以及熵感知调度调节注入强度。大量实验表明该框架从根本上消除了过度条件化伪影,在2步配置下实现了SOTA生成保真度。

英文摘要

While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (CFG) trajectory has emerged as the prevalent dual-dimensional compression paradigm. However, existing frameworks remain subjugated by a coarse-grained blind injection paradigm that perpetually enforces a globally static guidance strength while indiscriminately sampling the supervisor timestep. This state-agnostic design completely disregards the intrinsic nature of image generation as a dynamic evolutionary process characterized by progressive entropy reduction, which not only restricts the performance boundary of few-step compression but also precipitates severe CFG over-conditioning artifacts. To transcend these limitations, we re-examine the distillation procedure through the theoretical lens of Information Theory, formally modeling it as a dynamic mutual information game constrained by the Information Bottleneck (IB) principle. Specifically, we dismantle traditional blind assumptions via a dual-track adaptive framework. To determine the injection target, we propose an instance-aware selection mechanism that transmutes the intractable KL divergence constraint into a zero-overhead closed-form solution predicated on the local vector field norm. To regulate the injection strength, we introduce an entropy-aware schedule that dynamically decays alongside the SNR, applying maximal thrust for initial structural anchoring before smoothly reverting to the natural manifold to refine micro-details. Extensive empirical evaluations corroborate that our framework fundamentally eradicates over-conditioning artifacts, shattering the performance ceiling to achieve SOTA generative fidelity under extremely stringent 2-step configurations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑