发表机构
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Anhui University(广东省人工智能与数字经济实验室(深圳); 安徽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有扩散增强方法的缺陷,提出零样本退化感知Retinex引导扩散框架DARD,经多基准实验及语义分割下游验证,其性能优于现有零样本基线,可提升图像增强后的mIoU。
AI 中文摘要
现有基于扩散的增强方法在低光照图像增强(LLIE)中具备强大生成能力,但要么依赖配对监督,要么在零样本设置中缺乏可靠场景约束,常导致结构不一致与色彩偏移。受传统Retinex模型启发,该模型提供可物理解释的先验,可作为可靠场景约束却难以应对真实场景中的混合退化,我们提出DARD——一种用于LLIE的零样本退化感知Retinex引导扩散框架。DARD首先通过测试时的退化感知Retinex分解,从退化输入中提取图像特异性物理先验,为零样本恢复提供可靠结构引导;随后通过时间步自适应频率融合策略将这些先验注入反向扩散,以平衡结构锚定与细节生成;最后引入具备物理一致性和基于对比语言图像预训练(CLIP)的语义引导的引导反向细化过程,以在采样期间抑制结构伪影和语义偏移。大量实验表明,DARD在失真和感知性能上表现强劲,在多个真实低光照基准上始终优于现有零样本基线。为进一步验证该方法对下游应用的实用价值,我们评估了其对语义分割的影响,实验显示,经DARD增强的图像在平均交并比(mIoU)上相比AGLLDiff实现了28.10%的相对提升。
英文摘要
Existing diffusion-based enhancement methods provide strong generative capability for low-light image enhancement (LLIE), yet they either rely on paired supervision or lack reliable scene constraints in zero-shot settings, often leading to structural inconsistency and color drift. Motivated by conventional Retinex models, which offer physically interpretable priors that can serve as reliable scene constraints yet struggle with mixed degradations in real-world scenarios, we propose DARD, a zero-shot Degradation-Aware Retinex-guided Diffusion framework for LLIE. DARD first extracts image-specific physical priors from the degraded input through a test-time degradation-aware Retinex decomposition, thereby providing reliable structural guidance for zero-shot restoration. It then injects these priors into reverse diffusion through a timestep-adaptive frequency fusion strategy to balance structural anchoring and detail generation. Finally, a guided reverse refinement process with physical consistency and Contrastive Language-Image Pre-training (CLIP)-based semantic guidance is introduced to suppress structural artifacts and semantic drift during sampling. Extensive experiments show that DARD achieves strong distortion and perceptual performance and consistently outperforms existing zero-shot baselines across multiple real-world low-light benchmarks. To further validate the practical utility of our method for downstream applications, we evaluated its impact on semantic segmentation. Experiments demonstrate that images enhanced by DARD achieve a 28.10% relative improvement in mIoU over AGLLDiff.