arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RFMSR:用于图像超分辨率的残差流匹配

RFMSR: Residual Flow Matching for Image Super-Resolution

Shuwei Huang, Tianyao Luo, Jicheng Liu, Pan Zhou

arXiv 2607.12753首次发表:更新:

发表机构

Huazhong University of Science and Technology; Wuhan University(华中科技大学; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对图像超分辨率,提出残差流匹配框架RFMSR,将源分布集中于LQ潜变量以减少传输距离并保留结构先验,采用两阶段训练策略,实验证明该方法相比SOTA实现了相当甚至更优的感知质量。

AI 中文摘要

图像超分辨率(ISR)借助扩散模型和流匹配取得了显著进展。基于文本到图像(T2I)的主流方法利用大规模基础模型作为生成先验,虽实现了令人印象深刻的感知质量,但模型规模庞大且训练成本高昂。近期基于流匹配的纯视觉方法有显著进步,然而它们采用从纯高斯先验到数据分布的标准流公式,丢弃了低质量(LQ)输入中已有的丰富结构信息。此外,现有的单步加速技术往往丧失了模型的多步推理能力。本文提出用于图像超分辨率的残差流匹配(RFMSR),这是一个纯视觉框架,将源分布集中在LQ潜变量上,减少传输距离并在整个流轨迹中保留结构先验。还引入了两阶段训练策略:第一阶段通过条件流匹配预训练速度场,第二阶段对单步预测应用端到端监督,同时保留所有时间步的速度损失,在不牺牲多步细化的情况下实现高质量单步生成。大量实验表明,RFMSR与现有最先进(SOTA)方法相比,实现了相当甚至更优的感知质量。

英文摘要

Image super-resolution (ISR) has witnessed remarkable progress with diffusion models and flow matching. The dominant text-to-image (T2I) based approaches leverage large-scale foundation models as generative priors, achieving impressive perceptual quality but at the cost of massive model sizes and prohibitive training expenses. Recent flow-matching-based vision-only approaches have made significant strides; however, they adopt standard flow formulations that transport from a pure Gaussian prior to the data distribution, discarding the rich structural information already present in the low-quality (LQ) input. Furthermore, existing single-step acceleration techniques often forfeit the model's multi-step inference capability. In this paper, we propose Residual Flow Matching for Image Super-Resolution (RFMSR), a vision-only framework that centers the source distribution at the LQ latent, reducing transport distance and preserving structural priors throughout the flow trajectory. We further introduce a two-phase training strategy: Phase I pretrains the velocity field via conditional flow matching, while Phase II applies end-to-end supervision to the single-step prediction while retaining the velocity loss across all timesteps, achieving high-quality single-step generation without sacrificing multi-step refinement. Extensive experiments demonstrate that RFMSR achieves comparable or even superior perceptual quality compared to state-of-the-art (SOTA) methods. The source code is available at https://github.com/Faze-Hsw/RFMSR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑