arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S3VD:用于视频去雨的自适应语义引导时空扫描

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu

arXiv 2609.21322首次发表:更新:

发表机构

Harbin Institute of Technology; Zhengzhou Advanced Research Institute, Harbin Institute of Technology(哈尔滨工业大学; 哈尔滨工业大学郑州研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

S3VD框架通过多尺度语义融合和时空扫描融合模块,利用DINOv2语义先验和双门控Mamba,显著提升视频去雨性能,平均PSNR提高0.84 dB。

AI 中文摘要

强降雨通过破坏高频细节和引入运动模糊,严重降低了户外视频的质量,并严重削弱了视觉任务的可靠性。近年来,状态空间模型(SSMs),特别是Mamba,凭借其线性复杂度和对长距离依赖的建模能力,已成为视觉任务的高效替代方案。然而,当面对雨天视频中较差的视觉表示时,Mamba在保持2D空间语义的完整性和建模3D时空相关性方面仍面临困难。为突破这些限制,我们提出了S3VD,一个用于视频去雨的语义引导时空扫描框架,包含两个关键创新:多尺度语义融合(MSSF)模块和时空扫描融合(STSF)模块。前者整合来自DINOv2的时间语义先验,以引导精确的特征表示,并抵消Mamba的1D展平操作固有的局部语义上下文丢失,增强对极端退化的鲁棒性。后者引入了一种时空扫描机制,并设计了解耦门控Mamba(DG-Mamba)层,该层采用两个独立的门控来自适应地控制输入片段中的前后上下文信息,优化帧内和帧间相关性建模。在视频去雨基准上的实验证明了S3VD的优越性,与基于Mamba的基线相比,平均PSNR提高了0.84 dB,达到了最先进的性能。

英文摘要

Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear complexity and ability to model long-range dependencies. However, when confronted with the poor visual representations in rainy videos, Mamba still faces difficulties in preserving the integrity of 2D spatial semantics and modeling 3D spatio-temporal correlations. To break these limitations, we introduce S3VD, a Semantic-Guidance Spatio-Temporal Scanning framework for video deraining, featuring two key innovations: Multi-Scale Semantic Fusion (MSSF) Module and Spatio-Temporal Scanning Fusion (STSF) Module. The former integrates temporal semantic priors from DINOv2 to guide precise feature representation and counteract the loss of local semantic context inherent to Mamba's 1D flatten operation, enhancing robustness against extreme degradation. The latter introduces a spatio-temporal scanning mechanism and devises a Decoupled-Gating Mamba (DG-Mamba) layer, which employs two independent gates to adaptively control preceding and subsequent contextual information within the input clip, optimizing intra-frame and inter-frame correlation modeling. Experiments on video deraining benchmarks demonstrate the superiority of S3VD, achieving state-of-the-art performance with an average 0.84 dB PSNR improvement over Mamba-based baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑