arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16870cs.CV

tcnerv:隐式神经视频压缩中的双域时间上下文建模

tcnerv:dual-domain temporal context modeling for implicit neural video compression

Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

TCNeRV通过特征域多尺度门控融合与嵌入域残差编码,在约3M参数下实现UVG数据集36.08dB的PSNR,显著优于现有方法。

中文摘要 AI 辅助

视频压缩旨在在受限比特率下最小化重建失真。现有的视频隐式神经表示(INRs)通常独立解码帧,使中间特征不受先前重建结果的条件约束,且内容嵌入缺乏显式的时间预测。我们提出TCNeRV,在特征域和嵌入域中均利用重建上下文。其多尺度时间上下文融合(MTCF)模块在多个解码器尺度注入门控历史特征,而时间嵌入残差编码(TERC)预测每个内容嵌入并仅编码其残差。TCNeRV拥有约300万参数,在UVG数据集上平均PSNR达36.08 dB,比HNeRV-Boost高出2.20 dB。相对于HM、DCVC和HiNeRV,其BD-rate分别降低22.06%、66.73%和29.85%,在有限模型容量下展示了有竞争力的率失真性能。

英文摘要

Video compression aims to minimize reconstruction distor tion under a constrained bit rate. Existing video implicit neural representations (INRs) often decode frames independently, leaving intermediate features unconditioned on previous reconstructions and content embeddings without explicit temporal prediction. We propose TCNeRV, which exploits reconstructed context in both feature and embedding domains. Its multi-scale temporal-context fusion (MTCF) module injects gated historical features at multiple decoder scales, while temporal embedding-residual coding (TERC) predicts each content embedding and codes only its residual. With approximately 3M parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV-Boost by 2.20 dB. It reduces BD-rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV, respectively, demonstrating competitive rate-distortion performance with limited model capacity.

发表机构

  • Harbin Engineering University(哈尔滨工程大学)
  • Key Laboratory of Advanced Marine Communication and Information Technology(先进海洋通信与信息技术重点实验室)
  • University of Toronto(多伦多大学)
  • Kanagawa University(神奈川大学)

机构由 AI 辅助整理,请以论文原文为准。

↑