发表机构
Huazhong University of Science and Technology; Mohamed bin Zayed University of Artificial Intelligence; National University of Defense Technology; Xiaomi Communications Company Ltd.(华中科技大学; 穆罕默德·本·扎耶德人工智能大学; 国防科技大学; 小米通讯有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于扩散先验的曝光校正框架DPEC,通过高效微调策略与联合交叉注意力模块,在多数据集上实现了优于现有方法的曝光校正性能。
AI 中文摘要
尽管大多数现有的曝光校正方法能实现高保真度,但它们往往过度关注整体像素级精度,难以有效建模极端曝光区域,导致感知质量欠佳。近来,扩散模型因在图像生成领域的卓越性能受到广泛关注,然而将其成功应用于曝光校正仍是一个具有挑战性的开放问题,核心难点在于在随机扩散过程中生成准确的图像结构并保持高图像保真度。本文提出DPEC(Diffusion Prior-based Exposure Correction),一种利用预训练大规模扩散模型中封装的基于扩散的图像生成先验的新型图像曝光校正框架。具体而言,我们首先提出一种高效微调策略,从预训练模型中导出曝光校正器,使其能在单步去噪过程中生成增强图像;此外,我们无缝结合扩散模型与回归模型的优势,设计联合交叉注意力模块以整合多尺度扩散先验特征,从而有效保留高频细节并最小化随机伪影。扩散模型专注于处理低频内容而非所有复杂纹理细节。实验结果表明,所提出的DPEC方法在多个曝光校正数据集上,无论是在保真度、感知质量还是视觉效果方面,均持续优于现有最先进的方法。
英文摘要
Although most existing exposure correction methods achieve high fidelity, they often place excessive focus on overall pixel-wise accuracy, making it challenging to effectively model extreme exposure regions, which results in suboptimal perceptual quality. Recently, diffusion models have received significant attention due to their remarkable performance in the realm of image generation. However, their successful application to exposure correction remains a challenging and open question. The key challenge lies in generating accurate image structures and maintaining high image fidelity during stochastic diffusion processes. In this paper, we propose DPEC (Diffusion Prior-based Exposure Correction), a novel framework for image exposure correction that utilizes diffusion-based image generation priors encapsulated in pre-trained large-scale diffusion models. Specifically, we first propose an efficient fine-tuning strategy to derive an exposure corrector from pre-trained models, enabling the generation of enhanced images in a single-step denoising process. Moreover, we seamlessly combine the strengths of diffusion models and regression models, and design a joint cross-attention module to integrate multi-scale diffusion prior features, thereby effectively preserving high-frequency details and minimizing random artifacts. The diffusion model focuses on dealing with low-frequency content rather than all the intricate texture details. The experimental results demonstrate that the proposed DPEC method consistently outperforms existing state-of-the-art methods on multiple exposure correction datasets, whether in terms of fidelity, perceptual quality, or visual effects.
CommentsAccepted by IEEE Transactions on Multimedia (TMM)