MagnifiQ:面向高分辨率图像修复的感知补丁文本引导渐进式上采样方法
MagnifiQ: Patch-aware Text Guided Progressive Upscaling for High-Resolution Image Restoration
浏览论文内容
中文总结 AI 辅助
MagnifiQ是一种跨分辨率渐进式上采样的图像修复框架,替换扩散模型自注意力层为线性计算卷积,采用补丁文本提示,在4K图像修复中性能优于现有方法,实现速度与质量的实用权衡。
中文摘要 AI 辅助
从退化输入中恢复高分辨率图像颇具挑战,因为它必须在恢复精细局部细节的同时保持全局结构一致性,尤其是在4K分辨率下,直接基于扩散模型的修复方法计算成本高昂,且容易产生重复或不一致的纹理。在本研究中,我们提出了MagnifiQ,这是一种图像修复框架,可跨分辨率逐步上采样并修复图像,例如从1024×1024到4096×4096。我们的方法利用预训练的文本到图像扩散模型(如SDXL),通过将其原始自注意力层替换为计算成本随图像分辨率线性增长的卷积操作,使其适配更具可扩展性的高分辨率推理。我们进一步提出了一种渐进式上采样策略,在多个分辨率阶段迭代修复图像,优化每个中间输出而非直接生成最终的4K图像,从而提升全局一致性并减少高分辨率伪影。为增强局部细节同时控制内容漂移,MagnifiQ采用了针对特定补丁的文本提示,在修复过程中提供空间定位的语义引导。在合成及真实世界退化图像上的大量实验表明,MagnifiQ在感知质量和人类偏好方面优于现有基于扩散模型的修复方法,生成更清晰的纹理和更一致的4K结果,同时通过其可扩展的骨干网络和渐进式设计实现了实用的速度-质量权衡。
英文摘要
High-resolution image restoration from degraded inputs is challenging because it must preserve global structural consistency while recovering fine-grained local details, especially at 4K resolution where direct diffusion-based restoration is computationally expensive and prone to repeated or inconsistent textures. In this work, we introduce MagnifiQ, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096. Our approach leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-attention layers with convolutional operations whose computational cost grows linearly with image resolution. We further propose a progressive upscaling strategy that iteratively restores images over multiple resolution stages, refining each intermediate output rather than directly hallucinating the final 4K image, thereby improving global coherence and reducing high-resolution artifacts. To enhance local details while controlling content drift, MagnifiQ uses patch-specific text prompts that provide spatially localized semantic guidance during restoration. Extensive experiments on synthetic and real-world degraded images show that MagnifiQ outperforms prior diffusion-based restoration methods in perceptual quality and human preference, producing sharper textures and more coherent 4K results while offering practical speed--quality trade-offs through its scalable backbone and progressive design.
发表机构
- Qualcomm AI Research(高通人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。