arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

URVC:一种具有时间、空间和感知适应性的统一实时神经视频编码模型

URVC: A Unified Real-Time Neural Video Coding Model with Temporal, Spatial, and Perceptual Adaptivity

Xihua Sheng, Chang Wen Chen

arXiv 2607.15033首次发表:更新:

AI 中文总结

针对现有神经视频编解码器在动态环境中适应性差的问题,提出URVC。通过速率感知自适应时间预测、基于分解的空间速率控制和感知切换方法,使编解码器具备时间、空间和感知适应性,满足不同视频内容、用户需求及质量偏好。

AI 中文摘要

神经视频编码发展迅速,在实现有竞争力的压缩性能的同时还能达到实时编码速度。然而,现有编解码器在动态环境中部署时表现出严重的刚性,无法适应不同视频内容、用户需求和质量偏好。具体表现为:为满足实时约束丢弃显式运动估计和压缩,空间比特分配策略粗糙且固定,无法适应动态用户需求,不能在不部署单独模型的情况下适应质量偏好。本文提出URVC,通过速率感知自适应时间预测方法、基于分解的空间速率控制方法和感知切换方法,将刚性系统转变为具有时间、空间和感知适应性的统一框架。

英文摘要

Neural video coding has advanced rapidly, achieving competitive compression performance while also enabling real-time coding speed. Yet, existing codecs exhibit severe rigidity when deployed in dynamic environments, failing to adapt to different video content, user requirements, and quality preferences. First, to meet the real-time constraint, they discard explicit motion estimation and motion compression, thereby losing the ability to adapt temporal prediction to motion complexity and bitrate constraints. Second, their spatial bit allocation strategy is coarse and, once trained, is fixed. It cannot adapt to dynamic user requirements at test time, preventing users from freely controlling the spatial distribution of bits. Third, they cannot adapt their quality preference to varying application requirements without deploying separate models. We address all three limitations within a single real-time neural video codec--URVC, transforming a rigid system into a unified framework with temporal, spatial, and perceptual adaptivity. First, we propose a rate-aware adaptive temporal prediction method that generates diverse prediction candidates through a multi-candidate architecture and couples candidate selection directly to rate-distortion optimization. Second, we propose a decomposition-based spatial rate control method that achieves finer-grained spatial bit allocation through feature decomposition and separate quantization, and allows users to perform direct spatial rate control at test time without retraining. Third, we propose a perceptual switching method that only requires learning a secondary module bank alongside a frame generator, enabling a codec to switch between signal fidelity and perceptual quality modes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑