P-PatchDiff:用于低光照图像增强的渐进式补丁扩散模型
P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement
浏览论文内容
中文总结 AI 辅助
针对现有补丁扩散模型无法捕捉全局亮度上下文或计算成本高的局限,提出P-PatchDiff框架,通过动态调整补丁尺寸和多补丁对齐策略,高效实现多尺度低光照图像增强,速度提升80倍且内存占用低。
中文摘要 AI 辅助
近期低光照图像增强领域的进展利用了扩散模型,因其具备生成感知上真实、细节丰富图像的强大能力。补丁扩散模型进一步为与尺寸无关的图像复原提供了有前景的解决方案,同时提升了效率。然而,现有方法通常依赖小型固定补丁(如64×64),无法捕捉图像级的亮度上下文;而扩大感受野虽能改善亮度和颜色估计,但会大幅增加计算成本。此外,低光照图像常存在区域间亮度不均的问题,因此需确保局部增强的补丁组合成完整图像时保持视觉连贯性。为解决这些局限,我们提出P-PatchDiff,一种用于低光照图像增强的可扩展渐进式补丁扩散框架,其在去噪过程中动态调整补丁尺寸,实现从局部到全局视角的逐步转变。我们还引入了多补丁对齐策略,利用估计的全局亮度代理归一化不同补丁尺度的特征。P-PatchDiff不追求像素级重建精度,而是聚焦于可扩展性和整幅图像的亮度连贯性,使模型能感知多尺度信息,更好地增强亮度各异的区域。我们通过实验证明,P-PatchDiff可有效增强400×600至4K分辨率的图像,比现有补丁扩散模型快80倍,且内存使用不足9GB。代码可在该https URL获取。
英文摘要
Recent advancements in low-light image enhancement have leveraged diffusion models for their strong ability to generate perceptually realistic, detailed images. Patch diffusion models further offer a promising solution to size-agnostic image restoration while improving efficiency. However, existing methods typically rely on small, fixed patches (e.g., 64$\times$64) that cannot capture image-level brightness context, whereas enlarging the receptive field improves brightness and colour estimation but substantially increases computational cost. Moreover, low-light images often exhibit uneven brightness across regions, making it necessary to ensure that locally enhanced patches remain visually coherent when combined into the full image. To address these limitations, we propose P-PatchDiff, a scalable progressive patch diffusion framework for low-light image enhancement that dynamically adjusts patch size throughout the denoising process, enabling a gradual shift from local to global views. A Multi-Patch Alignment strategy is also introduced to normalise features across varying patch scales using an estimated global brightness proxy. Rather than pursuing pixel-level reconstruction accuracy, P-PatchDiff focuses on scalability and coherent brightness across the whole image, allowing the model to perceive multi-scale information and better enhance regions with varying brightness. We empirically demonstrate that P-PatchDiff effectively enhances images ranging from 400 $\times$ 600 to 4K and is 80$\times$ faster than existing patch diffusion models while using less than 9GB of memory. The code is available at https://github.com/RuoyuGuo/P-PatchDiff.
发表机构
- School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。