AI 中文总结
针对野外视频阴影去除难题,提出WildShadowRemover框架,通过LoRA微调适配视频扩散模型,结合细节注入、频率分解调制及Depth Anything 3深度先验,构建对应数据集,在阴影去除质量与时间一致性上优于现有方法。
AI 中文摘要
野外视频阴影去除因复杂光照、多样阴影外观及训练数据有限而颇具挑战性,尽管其对众多视觉与图形应用至关重要,但在无约束真实场景中仍未得到充分探索。为解决这一缺口,我们提出WildShadowRemover框架,该框架通过LoRA微调适配预训练视频扩散模型以实现鲁棒视频阴影去除。为在保留模型强大生成先验的同时保留图像细节,我们为冻结的VAE解码器增设细节注入模块,并引入阴影掩码引导的频率分解调制模块,以选择性恢复高频纹理并抑制阴影伪影。来自Depth Anything 3的单目深度先验进一步为具有挑战性的光照条件提供几何感知引导。我们还构建了WildShadow这一大型配对视频阴影去除数据集与基准,涵盖多样合成场景。大量实验表明,我们的方法在阴影去除质量和时间一致性方面优于现有方法,生成时间连贯的无阴影视频,具有卓越视觉质量且在具有挑战性的野外场景中表现出强泛化能力。
英文摘要
Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.