arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00483eess.IVcs.MM

可流式神经视频压缩:面向跨平台部署的混合精度方法

Streamable Neural Video Compression: A Mixed Precision Approach for Cross-Platform Deployment

Kasidis Arunruangsirilert, Heming Sun, Jiro Katto

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对神经视频编解码器跨平台部署的浮点非确定性问题,提出混合精度策略,实现1080p下GPU跨代无缝解码且压缩效率损失可忽略,验证了可流式学习视频压缩的可行性。

中文摘要 AI 辅助

神经视频编解码器(Neural Video Codecs, NVC)具备前所未有的率失真性能,在5G蜂窝网络、新兴卫星直连小区(Direct-to-Cell, D2C)链路等带宽受限环境中极具应用潜力。然而,NVC在实际流式应用部署中面临跨平台浮点非确定性问题,会导致算术熵编码器在不同GPU架构间失步崩溃。近期基于整数的量化方法虽能缓解该问题,但INT8会造成压缩效率大幅下降,INT16则因绕过硬件加速引发严重计算瓶颈。本文提出一种可流式的客户端-服务器NVC架构,采用新型混合精度(FP16/FP32)策略:在硬件加速的FP16中策略性执行P帧以保障实时吞吐量,同时将I帧及周期性特征适配器重置强制为符合IEEE-754标准的FP32,确保关键边界的确定性同步。通过在涵盖四个架构代际的12款GPU上开展广泛的编解码交叉评估,证明该方法成功消除代内碎片化,大幅提升跨芯片互操作性,实现1080p分辨率下近期架构间的无缝跨代解码,且对压缩效率影响可忽略不计。此外,本文在Wi-Fi 6、5G NR(FDD/TDD)、星链D2C等多种实际网络中评估系统端到端延迟,验证了可流式学习视频压缩的实际可行性,同时凸显了非地面网络中的独特挑战。

英文摘要

Neural Video Codecs (NVCs) offer unprecedented rate-distortion performance, making them highly attractive for bandwidth-constrained environments like 5G cellular networks and emerging satellite direct-to-cell (D2C) links. However, deploying NVCs in real-world streaming applications is severely hindered by cross-platform floating-point non-determinism, which causes arithmetic entropy coders to desynchronize and crash across different GPU architectures. While recent integer-based quantization methods address this, they incur either massive degradation in compression efficiency (INT8) or severe computational bottlenecks by bypassing hardware acceleration (INT16). In this paper, we propose a streamable, client-server NVC architecture featuring a novel Mixed Precision (FP16/FP32) strategy. By strategically executing P-frames in hardware-accelerated FP16 for real-time throughput, while forcing I-frames and periodic feature-adapter resets to IEEE-754 compliant FP32, we guarantee deterministic synchronization at critical boundaries. Through extensive cross-encode/decode evaluations across 12 GPUs spanning four architectural generations, we demonstrate that our approach successfully eliminates intra-generation fragmentation and substantially broadens cross-die interoperability, achieving seamless cross-generation decodability for recent architectures at 1080p. Crucially, this is achieved with a negligible impact on compression efficiency. Furthermore, we evaluate the system's end-to-end latency across diverse real-world networks, including Wi-Fi 6, 5G NR (FDD/TDD), and Starlink D2C, proving the practical viability of streamable learned video compression while highlighting unique challenges in Non-Terrestrial Networks.

补充信息

↑