发表机构
School of Computer Science and Technology, Beijing Institute of Technology; Chinese Aeronautical Establishment; Westlake University; Shanghai Jiao Tong University(北京理工大学计算机科学与技术学院; 中国航空研究院; 西湖大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出二值化RAW视频恢复框架BinRVR,通过BIIM和DAB-Conv实现96%计算量与参数缩减,仅4%性能下降,在多项RAW视频恢复任务上表现具竞争力,还可拓展至下游视频应用。
AI 中文摘要
RAW视频恢复是高质量低层视觉感知的基础,也是众多下游视觉应用的前提。二值神经网络(BNNs)虽能实现图像增强的高效轻量部署,但其在时间一致性和激活值分布建模上的不足,限制了在视频场景中的应用。本文提出BinRVR,一种二值化RAW视频恢复框架,可减少约96%的计算量和参数,仅带来约4%的性能下降。具体而言,我们提出二值化信息交互模块(BIIM),以高效统一的方式联合建模空间与时间信息;还开发了分布感知二值化卷积(DAB-Conv),利用全精度激活的统计信息缓解量化误差。该框架还支持多位量化,可在不同硬件约束下灵活权衡精度与效率。大量实验表明,BinRVR在RAW视频恢复任务(包括低光增强、去噪、去模糊、超分辨率)上,与现有最优二值化方法相比具备竞争力;我们还进一步探索了该方法在下游视频应用(目标检测、单目深度估计)中的潜力。
英文摘要
RAW video restoration is fundamental to high-quality low-level perception and serves as the basis for a wide range of downstream vision applications. While binary neural networks (BNNs) enable efficient lightweight deployment for image enhancement, their deficiencies in modeling temporal coherence and activation value distributions hinder their effectiveness when applied to video scenarios. In this paper, we propose BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation. Specifically, we present a Binarized Information Interaction Module (BIIM) to jointly model spatial and temporal information in an efficient and unified manner. Moreover, we develop a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors. The proposed framework further supports multi-bit quantization, enabling flexible accuracy-efficiency trade-offs across different hardware constraints. Extensive experiments demonstrate that our BinRVR achieves competitive performance compared with state-of-the-art binarized methods on RAW video restoration tasks, including low-light enhancement, denoising, deblurring, and super-resolution. We further explore the potential of our method on downstream video applications, including object detection and monocular depth estimation.
CommentsAccepted by TPAMI2026