arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09133cs.CVcs.AI

当潜变量遗忘像素时:在扩散Transformer超分辨率中恢复保真度

When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution

Yu Shi, Yuyao Zhang, Yu-wing Tai

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对扩散Transformer超分辨率中VAE压缩导致的保真度下降问题,提出像素锚定超分辨率框架,通过复用VAE前的像素证据提升结果的忠实度,实验验证其性能优于现有潜变量生成式超分辨率方法。

中文摘要 AI 辅助

基于大型生成模型的图像超分辨率(SR)近期已实现出色的感知质量,但保持与低分辨率(LR)观测的保真度仍具挑战性。尤其,我们发现基于潜变量表示的扩散Transformer(DiT)存在关键局限:变分自编码器(VAE)的压缩瓶颈削弱了细粒度空间信息,导致与输入图像关联度弱的幻觉细节。本研究从表示视角重新审视生成式SR,提出像素锚定超分辨率(PGSR)框架,该框架在VAE压缩前保留LR观测的像素证据并在整个恢复过程中复用。PGSR不依赖压缩后的潜变量条件,而是从上采样LR图像中提取VAE前的像素证据并在两个阶段复用:第一,条件侧轨迹引导将LR衍生的像素证据与潜变量LR条件融合,以引导潜变量恢复轨迹;第二,解码器侧像素锚定将多尺度像素特征注入冻结的VAE解码器,利用LR观测线索锚定最终渲染结果。为高效适配大型预训练DiT模型,我们保持潜变量自编码器和主流匹配骨干冻结,仅训练轻量恢复模块。我们还研究了一种高效的局部窗口注意力变体,以提升高分辨率效率与可扩展性。大量实验表明,PGSR改善了真实感-保真度的权衡,相比现有潜变量生成式SR方法生成更忠实、视觉更具说服力的结果。

英文摘要

Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformers (DiTs) built on latent representations suffer from a critical limitation: the compression bottleneck of the VAE weakens fine-grained spatial information, leading to hallucinated details that are weakly grounded in the input image. In this work, we revisit generative SR from a representation perspective and propose a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration. Instead of relying solely on the compressed latent condition, PGSR extracts pre-VAE pixel evidence from the upsampled LR image and reuses it at two stages. First, Condition-Side Trajectory Guidance fuses LR-derived pixel evidence with the latent LR condition to guide the latent restoration trajectory. Second, Decoder-Side Pixel Grounding injects multi-scale pixel features into the frozen VAE decoder to ground the final rendering with LR-observed cues. To efficiently adapt large pretrained DiT models, we keep the latent autoencoder and main flow-matching backbone frozen, and train only lightweight restoration modules. We further study an efficient local-window attention variant for improved high-resolution efficiency and scalability. Extensive experiments demonstrate that PGSR improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.

发表机构

  • Dartmouth College(达特茅斯学院)

机构由 AI 辅助整理,请以论文原文为准。

↑