发表机构
University of California, Berkeley; Alibaba Group; University of Michigan, Ann Arbor; Nanjing University; Southeast University(加州大学伯克利分校; 阿里巴巴集团; 密歇根大学安娜堡分校; 南京大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在冻结的DINOv3-L潜在空间中进行超分辨率,该基底使退化嵌入靠近流形并含语义层级,训练415M解码器(RAESR)实现最佳保真度-感知权衡,优于现有方法。
AI 中文摘要
在图像潜在空间中,高分辨率、自然且清晰的图像的嵌入构成一个流形。高分辨率图像的退化会将其嵌入推离该流形。真实世界超分辨率(SR)因此成为将退化嵌入映射回该流形的任务——不是映射到流形上的任意位置,而是映射到保留输入仍携带的信息(包括语义和像素细节)的点。所有已发表的方法都在面向重建的潜在空间或像素空间中实现此映射。我们声称这些空间是SR的错误基底。低分辨率和退化图像嵌入远离流形,使得映射困难且代价高昂。这些基底中缺乏语义信息,也使得难以导航到流形上的忠实点,在退化严重时导致严重的幻觉。因此,在合适的潜在空间中恢复对SR任务至关重要。我们表明,冻结的DINOv3-L的23个融合层的潜在空间就是这样一个使SR任务更易处理的空间。退化图像嵌入在流形附近。此外,该基底包含从像素记录到退化鲁棒语义的信息层级,引导模型找到流形上的忠实点。在此基底上,一个415M解码器在重建和对抗目标下训练,以将退化嵌入映射回并以单次通过解码到像素空间。由此产生的模型RAESR在RealSR、DRealSR、LSDIR和DIV2K-Val上,在保真度-感知权衡方面达到了最先进的对抗和基于扩散的恢复器中的最佳性能,在单个H20 GPU上每张512×512图像耗时37毫秒。在相同配方下将基底替换为VAE潜在空间,在所有指标上均表现更差。
英文摘要
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.