能量引导的流匹配(Energy-Guided Flow Matching)
Energy-Guided Flow Matching
AI总结:
本文提出能量引导的流匹配(EG-FM),通过移动端点显式建模从粗到细的生成轨迹,无需额外适配,在ImageNet类条件图像生成及文本到图像生成任务中均取得优异性能。
AI中文摘要:
像素空间生成模型可避免有损的潜在压缩,但需在高维空间中联合学习全局结构与细粒度细节。标准流匹配将噪声插值到固定的清晰图像端点,频谱演化需隐式学习。本文提出能量引导的流匹配(EG-FM),通过移动端点显式建模从粗到细的生成轨迹;具体而言,EG-FM用热核滤波端点替代固定端点,该端点从低频图像平滑演化到清晰图像,且通过图像特定的能量引导调度释放移动端点中的高频信号占比,从而重新调整流中的速度。该框架无需适配骨干网络和训练数据,在训练与推理阶段带来的开销可忽略不计。实验中,EG-FM在ImageNet类条件图像生成任务(256×256分辨率)上以更少的轮次持续实现更低的FID,200轮时FID达1.55,600轮时达1.45;在512×512分辨率设置下继续训练该生成任务,仅经40轮高分辨率适配后FID达1.58;此外,将EG-FM迁移到文本到图像生成任务,在GenEval得分上达0.85,在DPG-Bench上达83.9,代码可在指定网址获取。
英文摘要:
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at $256 \times 256$ with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of $512 \times 512$ resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.