arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ODDR:通过奖励引导的一步式阴影去除扩散模型

ODDR: One-Step Deshadow Diffusion via Reward Guidance

Junseong Shin, Kijun Kim, Minseong Kim, Dongjin Kim, Tae Hyun Kim

arXiv 2610.01291首次发表:更新:

发表机构

Hanyang University; AX Future Technology Institute(汉阳大学; AX未来技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对阴影去除依赖昂贵成对数据的问题,提出ODDR框架,利用合成数据训练基线模型并引入无标注奖励模型ShadowReward进行引导微调,实现高效高保真的一步式阴影去除,缩小与全监督方法的差距。

AI 中文摘要

近年来,深度学习在阴影去除领域取得了显著进展,大幅提升了图像质量和真实感。然而,大多数方法依赖于真实世界的成对数据集,这些数据集的收集成本高昂,且场景多样性往往有限,导致泛化能力受限。为解决这些局限,我们提出了通过奖励引导的一步式阴影去除扩散模型(ODDR),这是一个新框架,无需依赖真实世界的成对监督即可实现高效且高保真的阴影去除。我们的方法始于一步式阴影去除扩散模型(ODD),这是一个在合成阴影数据上训练的基线模型,用于高效的一步式无阴影重建。我们进一步利用ShadowReward将ODD改进为ODDR。与传统依赖大量标注的方法不同,ShadowReward是首个完全无需人工标注训练的阴影去除奖励模型。它通过排序带有受控退化(如纹理失真和边界伪影)的合成生成图像,来学习模拟人类的感知判断。这种奖励引导的微调使ODDR能够缩小合成数据与真实数据之间的领域差距。大量实验表明,ODD在无需真实世界成对监督的情况下实现了强劲性能,而ODDR进一步提升了结果,缩小了与在真实世界成对数据上训练的完全监督方法的差距,同时作为单步模型保持了更高的计算效率。

英文摘要

Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑