arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28760cs.CVcs.AIcs.LGstat.ML

面向信号的WaiT:简单的频率感知流匹配

WaiT for the Signal: Simple Frequency-Aware Flow-Matching

Krunoslav Lehman Pavasovic, Théophane Vallaeys, Stéphane Mallat, Giulio Biroli, Luke Zettlemoyer, Brian Karrer, Jakob Verbeek

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出频率感知流匹配模型WaiT,通过无损小波分解生成过程,引入三轴评估协议,在ImageNet等数据集上实现SOTA生成性能,采样计算量降低50%,还可无缝扩展至视频生成。

中文摘要 AI 辅助

随着图像生成模型扩展到越来越高的分辨率,全局一致性、局部细节和纹理保真度成为生成质量的关键维度。然而,标准流匹配方法对所有空间频率一视同仁,忽略了自然的频率层级——高频带远比粗结构更早与纯噪声无法区分。我们引入WaiT,一种小波感知图像Transformer,它通过无损小波将生成过程分解为粗带和细带。正如其名,高频带等待信号出现:在粗结构形成前保持纯噪声状态,待粗结构出现后再加入流匹配进行联合优化。由于标准FID通过激进下采样丢弃了细粒度细节,我们引入更严格的三轴评估协议,以评估原生分辨率下的质量。在ImageNet 512×512数据集上,WaiT实现了像素空间FID为1.43,且在三个维度上均为帕累托最优,采样计算量最多降低50%。我们最大的2B模型,在ImageNet 512分辨率的像素空间模型上达到了1.3的新SOTA FID。我们的方法在纹理保真度上甚至优于最强的潜在空间模型,且可无缝扩展到高分辨率OpenImages和视频生成,在Kinetics-600上实现了0.84的SOTA FVD,且无需算法修改。

英文摘要

As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse structures. We introduce WaiT, a Wavelet-aware image Transformer that decomposes generation into coarse and fine bands via lossless wavelets. True to its name, the high-frequency bands wait for the signal: staying pure noise until coarse structure has emerged, then joining the flow for joint refinement. Since standard FID discards fine-grained detail through aggressive downsampling, we introduce a more stringent three-axis evaluation protocol to assess quality at native resolution. On ImageNet 512x512, WaiT achieves a pixel-space FID of 1.43 and is Pareto-optimal across all three axes, reducing sampling compute by up to 50%. With our largest 2B model, we set a new state-of-the-art FID of 1.3 for pixel-space models on ImageNet 512 resolution. Our formulation outperforms even the strongest latent-space models on texture fidelity, and scales seamlessly to high-resolution OpenImages and to video generation, achieving a state-of-the-art FVD of 0.84 on Kinetics-600 with no algorithmic modifications.

发表机构

  • FAIR, Meta(FAIR(Meta旗下研究机构))
  • École Normale Supérieure(巴黎高等师范学院)
  • Sorbonne University(索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑