arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于深度学习的视频编码滤波:架构、算法与复杂度分析综述

Deep Learning-based Filtering for Video Coding: A Survey on Architectures, Algorithms, and Complexity Analysis

Young-Woon Lee, Byung-Gyu Kim

arXiv 2607.16319首次发表:更新:

发表机构

Sunmoon University; Sookmyung Women’s University; ICT Convergence Research Institute(顺天乡大学; 淑明女子大学; 信息通信融合研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对超高清显示等带来的高效视频编码需求,对基于深度学习的视频编码滤波技术进行综述。提出三维分类法,分析率失真性能与硬件可行性权衡,纳入最新标准化活动,识别相关挑战,为下一代智能视频编码提供指导和路线图。

AI 中文摘要

随着超高清(UHD)显示和沉浸式媒体服务在物联网(IoT)和消费电子(CE)领域(包括8K显示和移动设备)变得无处不在,对高效视频编码的需求空前高涨。基于深度学习的滤波(DLF)虽有望减轻诸如高效视频编码(HEVC/H.265)和通用视频编码(VVC/H.266)等标准中固有的压缩伪像,但其在CE设备中的部署受计算复杂度、内存带宽和功耗严重限制。为弥合学术研究与实际部署间的差距,本文对DLF技术进行全面的、面向硬件的综述。提出系统的三维分类法,将方法分为视频编码内的集成方案、编码信息利用和网络设计策略。与先前综述不同,本文批判性地分析了率失真(RD)性能与硬件可行性之间的权衡,强调从重型、面向性能的模型向针对神经处理单元(NPUs)的轻量级、硬件友好型架构的演变。此外,纳入联合视频专家组(JVET)在基于神经网络的视频编码(NNVC)方面的最新标准化活动以提供实际指导方针。还识别了诸如实时推理延迟和错误传播等开放挑战,为下一代CE视觉端点中强大、低功耗的智能视频编码提供路线图。

英文摘要

As Ultra-High-Definition (UHD) displays and immersive media services become ubiquitous in the Internet of Things (IoT) and Consumer Electronics (CE) sectors, including 8K display and mobile devices, the demand for high-efficiency video coding is unprecedented. While Deep Learning-based Filtering (DLF) has emerged as a promising solution to mitigate compression artifacts inherent in standards like High Efficiency Video Coding (HEVC/H.265) and Versatile Video Coding (VVC/H.266), its deployment in CE devices is severely constrained by computational complexity, memory bandwidth, and power consumption. To bridge the gap between academic research and practical deployment, this paper presents a comprehensive, hardware-oriented survey of DLF techniques. We propose a systematic three-dimensional taxonomy classifying methods into (1) Integration Scheme within the Video Coding, (2) Coding Information Utilization, and (3) Network Design Strategy. Unlike prior reviews, this work critically analyzes the trade-offs between Rate-Distortion (RD) performance and hardware feasibility, highlighting the evolution from heavy, performance-oriented models to lightweight, hardware-friendly architectures targeting Neural Processing Units (NPUs). Furthermore, we incorporate the latest standardization activities from the Joint Video Experts Team (JVET) on Neural Network-based Video Coding (NNVC) to provide realistic guidelines. We also identify open challenges such as real-time inference latency and error propagation, providing a roadmap toward robust, low-power intelligent video coding in next-generation CE vision endpoints.

DOI:10.1109/TCE.2026.3711657

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑