发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Fill2SR通过重利用掩码修复扩散Transformer实现真实世界超分辨率,无需额外空间分支,支持混合分辨率训练,在合成和真实基准上取得优异性能。
AI 中文摘要
近期真实世界图像超分辨率(SR)方法通常采用文本到图像(T2I)骨干网络,并添加ControlNet风格分支或空间条件令牌,这增加了内存和计算量,且随分辨率增长,并常将训练限制在固定尺度。我们提出Fill2SR,该方法重新利用掩码修复扩散Transformer进行SR,无需额外空间分支。我们的修复接口证据适配器(IIEA)在整图掩码下将低质量(LQ)观测写入原生掩码图像槽位,将修复转化为反向退化条件整流流,仅通过LoRA微调进行训练。我们进一步引入RCDT,一种离线流程,从非配对真实图像中提取退化描述符,并使用冻结的开源模型将其转移到干净目标上。Fill2SR支持高达QHD的混合分辨率训练,并在$512/1024/2048$输出上表现稳定。在合成基准上,我们的带有IIEA的基础模型在DIV2K和LSDIR上取得最佳LPIPS;添加RCDT以轻微LPIPS下降换取在RealLQ250和RealPhoto60上更一致的无参考质量提升。Fill2SR保持内存可预测性,在单个32GB GPU上运行$1536^2$推理,并通过分块恢复扩展到多百万像素输出。
英文摘要
Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a masked-inpainting Diffusion Transformer for SR without extra spatial branches. Our Inpainting-Interface Evidence Adapter (IIEA) writes the low-quality (LQ) observation into the native masked-image slot under a full-image mask, turning inpainting into a reverse-degradation conditional rectified flow trained with LoRA-only tuning. We further introduce RCDT, an offline pipeline that distills degradation descriptors from unpaired real images and transfers them onto clean targets using frozen open-source models. Fill2SR supports mixed-resolution training up to QHD and yields stable performance across $512/1024/2048$ outputs. On synthetic benchmarks, our base model with IIEA achieves the best LPIPS on DIV2K and LSDIR; adding RCDT trades a small LPIPS drop for consistently stronger no-reference quality on RealLQ250 and RealPhoto60. Fill2SR remains memory-predictable, running $1536^2$ inference on a single 32GB GPU and extending to multi-megapixel outputs via tiled restoration.
CommentsECCV 2026. 26 pages, 12 figures, including an 8-page appendix with additional visual results
Journal refComputer Vision - ECCV 2026, Part XXIV, Lecture Notes in Computer Science, vol. 17024, pp. 441-457 (2026)
DOI:10.1007/978-3-032-37556-8_24