arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16279eess.IVcs.ITcs.LGcs.MMmath.IT

语义感知的神经视频编解码器,用于抗差错的低延迟传输

Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

  • Nokia Bell Labs(诺基亚贝尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan

中文总结 AI 辅助

针对不可靠信道下的低延迟视频传输,提出基于DCVC-RT的语义感知多级神经编码方法,通过优先级数据包分配和抗差错熵模型,显著提升鲁棒性并保留任务相关视觉内容。

中文摘要 AI 辅助

新兴的物理人工智能系统要求在不可靠的信道上进行低延迟、面向任务的视频通信。我们提出了一种语义感知的多级神经视频编码方法,用于在抽象为多级数据包擦除信道的不可靠信道上进行稳健的低延迟视频传输。该框架基于实时DCVC-RT神经视频编解码器,引入了一种语义和特征感知的编码策略,将编码表示划分为携带不同级别语义和潜在特征重要性的数据包,并将这些数据包分配给不同的流,每个流在不可靠通信信道上传输时都关联一个优先级。我们还开发了一种抗差错的熵模型,消除了数据包间的依赖关系,使得每个数据包在数据包丢失时能够独立解码。整个系统在抽象的多级数据包擦除信道上进行端到端训练,从而能够学习信道感知的表示以及重要性感知的数据包分配,同时促进网络进行差异化的数据包优先级排序。实验表明,所提出的框架在数据包擦除情况下显著提高了相对于基线DCVC-RT的稳健性,在较不重要的区域实现优雅降级,同时更好地保留与任务相关的视觉内容。

英文摘要

Emerging physical AI systems require low-latency, task-oriented video communication over unreliable channels. We propose a semantic-aware multi-level neural video coding method for robust low-latency video transmission over unreliable channels that are abstracted as multi-level packet erasure channels. Built upon the real-time DCVC-RT neural video codec, the proposed framework introduces a semantic- and feature-aware coding strategy that partitions encoded representations into packets carrying different levels of semantic and latent-feature importance and assigns these packets to different streams, each associated with a priority level when transmitted over unreliable communication channels. We also developed an error-resilient entropy model that removes inter-packet dependencies, allowing each packet to be decoded independently under packet losses. The complete system is trained end-to-end over the abstracted multi-level packet erasure channels, enabling learning of channel-aware representations together with importance-aware packet assignment while facilitating the network for differentiated packet prioritization. Experiments show that the proposed framework significantly improves robustness over baseline DCVC-RT under packet erasures, achieving graceful degradation in less important regions while better preserving task-relevant visual content.

↑