发表机构
King Fahd University of Petroleum and Minerals; Ontario Tech University; University of Western Australia; University of Canberra(法赫德国王石油与矿产大学; 安大略理工大学; 西澳大利亚大学; 堪培拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于CLIP引导提示学习的无监督统一背光与暗光图像增强框架,通过对称残差U-Net结合空洞空间金字塔池化模块实现自适应校正,在多数据集上优于现有方法,为图像增强提供鲁棒可扩展方案。
AI 中文摘要
背光图像和暗光图像常存在严重的曝光失衡或全局欠曝光问题,对视觉感知及下游计算机视觉任务构成重大挑战。本文提出一种统一的无监督增强框架,可同时应对两类退化问题,且无需依赖配对的真实数据。该方法基于CLIP引导的提示学习,利用学习到的正、负文本提示对增强过程进行语义监督;为提升性能,设计了对称残差U-Net骨干网络,并添加了空洞空间金字塔池化模块,该架构可捕获多尺度上下文信息,支持空间异质光照下的自适应校正。训练过程中,增强网络受CLIP引导的语义相似性损失约束,并通过迭代提示优化机制进行优化。在BAID、Backlit300、LOL、VE-LOL-L等配对和非配对数据集上开展的大量实验表明,本框架在保真度、感知质量和泛化性方面,均持续优于现有最先进的监督与无监督方法。此外,本研究强调需建立更严格的背光增强基准协议,该领域目前研究较少;所提框架为不同光照条件下的真实场景光照增强提供了一种鲁棒、可扩展的解决方案。
英文摘要
Backlit and low-light images often suffer from severe exposure imbalance or global underexposure, presenting significant challenges for both visual perception and downstream computer vision tasks. In this paper, we propose a unified, unsupervised enhancement framework that addresses both types of degradation without relying on paired ground-truth data. Our approach builds on CLIP-guided prompt learning to semantically supervise enhancement using learned positive and negative textual prompts. To improve the quality of our improvements over prior work, we design a symmetric residual U-Net backbone augmented with an Atrous Spatial Pyramid Pooling module. This architecture captures multi-scale contextual information, enabling adaptive correction under spatially heterogeneous illumination. During training, the enhancement network is guided by CLIP-based semantic similarity losses and refined via an iterative prompt optimization mechanism. Extensive experiments on both paired and unpaired datasets, including BAID, Backlit300, LOL, and VE-LOL-L, demonstrate that our framework consistently outperforms state-of-the-art supervised and unsupervised methods in terms of fidelity, perceptual quality, and generalization. Furthermore, our work emphasizes the need for stronger benchmarking protocols for backlit enhancement, a relatively underexplored area. The proposed framework provides a robust, scalable solution for real-world illumination enhancement across diverse lighting conditions.