arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25275cs.CVcs.AI

ScaleResfusion:基于残差向量场的残差整流流

ScaleResfusion: Residual Rectified Flow based on Residual Vector Field

  • Nankai University(南开大学)
  • University of Macau(澳门大学)
  • Csiro Data 61(联邦科学与工业研究组织数据61部)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhenning Shi, Chen Xu, Junhao Zhang, Kefei Zhang, Linjie Liu, Zhedong Zheng, Tao Li

AI总结:

研究旨在解决现实世界图像恢复问题,提出ScaleResfusion框架,核心是残差整流流,从低质量图像出发学习残差向量场,结合知识蒸馏降低采样成本,实验证明其高效且性能优,为图像恢复提供实用可扩展方法。

AI中文摘要:

现实世界图像恢复旨在从复杂且未知的退化中恢复高质量图像。近期基于扩散的方法虽提升了感知质量,但仍有两个关键挑战未解决。本文提出ScaleResfusion,一种基于预训练文本到图像整流流模型的可扩展扩散框架。核心是残差整流流,引入残差项R,从有噪声的低质量图像开始,通过学习残差向量场,使输出分布等与预训练模型一致,还引入知识蒸馏降低采样成本。实验表明该方法高效且性能达最优,为适配大预训练扩散模型到现实图像恢复提供实用可扩展方式。

英文摘要:

Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Recent diffusion-based methods have substantially improved perceptual quality, yet two obstacles remain: methods that sample from Gaussian noise require many steps and are often less faithful to the degraded input, whereas residual-based methods that start from the low-quality (LQ) image typically train task-specific models from scratch, with optimization objectives coupled to a particular noise scheduler, and therefore cannot reuse modern pre-trained generative priors. We present \textbf{ScaleResfusion}, which rewrites residual restoration as a scheduler-independent adaptation interface for pre-trained text-to-image rectified-flow models. Its core, \textbf{Residual Rectified Flow} (RRF), inserts the residual term $R$ into the linear transport path of Rectified Flow, so that sampling starts from noisy LQ at an exact acceleration point, where the signal-to-noise ratio of the starting state is continuously controlled by the residual ratio $γ$. The resulting optimization target, the \textbf{residual vector field}, contains no scheduler-specific coefficients and differs from the pre-trained rectified-flow target only by the residual offset $γR$; adapting a frozen billion-scale backbone therefore reduces to fitting this compact residual correction with LoRA-only training. A knowledge-distillation pipeline built around RRF further reduces sampling to as few as 4 steps. Experiments on real-world super-resolution across multiple benchmarks show that ScaleResfusion achieves state-of-the-art restoration quality and transfers consistently across pre-trained rectified-flow backbones from 2B to 9B parameters.

↑