发表机构
School of Computer Science and Technology, Harbin Institute of Technology; School of Computer Science and Technology, Harbin University of Science and Technology(哈尔滨工业大学计算机科学与技术学院; 哈尔滨理工大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频去雨问题,提出RainDancer框架,基于分解-交互范式,在RGB和事件分支分别处理,分离雨和背景成分,进行组件级融合,并引入事件域监督,实验表明该方法在多方面性能优越。
AI 中文摘要
视频去雨旨在从雨天视频中恢复干净的视觉内容,以便在恶劣天气下进行可靠感知。现有方法主要依赖RGB序列和时间冗余,但仅RGB恢复在动态雨景中仍不明确。事件相机提供具有高时间分辨率的互补运动敏感线索,但事件流也包含传感器噪声和背景触发响应。为解决此问题,我们提出RainDancer,一种基于分解-交互范式的渐进式RGB-事件视频去雨框架。核心思想是在跨模态交互之前在每个模态内分离雨和背景成分。在RGB分支中,帧特征被逐步分解为雨和背景表示。在事件分支中,一个面向降雨的脉冲神经网络模块捕获与雨运动相关的稀疏和突发事件动态。然后在语义对齐的表示之间进行组件级融合以保持结构和抑制降雨。我们还引入事件域监督来规范稀疏事件重建、结构一致性和梯度方向。在合成和真实RGB-事件视频去雨数据集上的实验证明了其优越的定量性能、视觉质量和下游感知鲁棒性。
英文摘要
Video deraining aims to recover clean visual content from rainy videos for reliable perception under adverse weather. Existing methods mainly rely on RGB sequences and temporal redundancy, but RGB-only restoration remains ambiguous in dynamic rainy scenes, where rain streaks, textures, boundaries, motion, and occlusions may share similar visual patterns. Event cameras provide complementary motion-sensitive cues with high temporal resolution, but event streams also contain sensor noise and background-triggered responses, so direct RGB-Event fusion may introduce cross-modal interference. To address this issue, we propose RainDancer, a progressive RGB-Event video deraining framework based on a decompose-before-interact paradigm. The core idea is to separate rain and background components within each modality before cross-modal interaction. In the RGB branch, frame features are progressively decomposed into rain and background representations. In the event branch, a rain-oriented spiking neural network module captures sparse and bursty event dynamics associated with rain motion. Component-level fusion is then performed between semantically aligned representations for structure preservation and rain suppression. We further introduce event-domain supervision to regularize sparse event reconstruction, structural consistency, and gradient orientation. Experiments on synthetic and real RGB-Event video deraining datasets demonstrate superior quantitative performance, visual quality, and downstream perception robustness. Code is available at https://github.com/AE86-plus/RainDancer.