XYZFlow:用于高效生成建模的多维捷径流缩放方法
XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling
查看机构详情
- CUHK(香港中文大学)
- Westlake University(西湖大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
XYZFlow框架通过流匹配的多维缩放实现高效图像生成,兼具7.2-8.5倍的教师模型速度提升与竞争力FID,其下一捷径预测可实现更优的质量-延迟权衡。
中文摘要 AI 辅助
高保真图像生成面临速度与质量的权衡问题。扩散模型能生成高质量图像,但需要代价高昂的迭代采样。现有高效方法主要是将预训练模型蒸馏为少步采样器,这一过程极具挑战性,且高度依赖教师模型的质量。本文提出XYZFlow框架,通过对流匹配进行多维缩放来重新思考高效生成。与单步映射不同,XYZFlow通过结构化多维条件使概率路径更具可识别性和可学习性,从而增强表达能力。我们将自回归建模视为隐式流拉直,更丰富的上下文可减少轨迹歧义。XYZFlow通过两个正交维度实现这一思路:时间缩放,对完整去噪历史使用非马尔可夫条件;空间缩放,通过下一捷径预测实现,利用前序图像块的去噪轨迹作为先验,按顺序生成图像块。实验表明,XYZFlow实现了最先进的性能,教师模型速度提升7.2至8.5倍,FID指标具有竞争力,且下一捷径预测相比模型缩放或减少步数,能实现更优的质量-延迟权衡。
英文摘要
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a challenging process that depends heavily on teacher-model quality. In this paper, we introduce XYZFlow, a framework that rethinks efficient generation through multidimensional scaling of flow matching. Unlike single-step mappings, XYZFlow enhances expressivity by making probability paths more identifiable and learnable through structured multidimensional conditioning. We view autoregressive modeling as implicit flow straightening, where richer context reduces trajectory ambiguity. XYZFlow realizes this idea through two orthogonal dimensions: temporal scaling, which uses non-Markovian conditioning on the full denoising history; and spatial scaling, enabled by Next Shortcut Prediction, which sequentially generates patches using preceding patches' denoising trajectories as priors. Experiments show that XYZFlow achieves state-of-the-art performance, with 7.2-8.5X teacher speedups and competitive FID, while Next Shortcut Prediction delivers superior quality-latency trade-offs over model scaling or step reduction.