arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向360度视频的投影感知端到端学习型视频压缩

Projection-Aware End-to-End Learned Video Compression for 360-Degree Video

Niloofar Maani

arXiv 2608.28689首次发表:更新:

发表机构

Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探讨投影格式对360度视频端到端神经压缩的影响,评估7种投影格式,发现投影效率依赖于编解码器,为学习型360度视频压缩的投影选择提供指导。

AI 中文摘要

360度视频支持虚拟现实、自动驾驶和教育等沉浸式应用。由于传统视频编解码器无法直接处理球形内容,必须先将其映射为二维投影。投影选择会影响空间连续性、采样均匀性、运动估计和压缩效率。本研究探讨投影格式如何影响360度视频的端到端神经压缩。使用scale-space flow模型、JVET测试序列和通用测试条件,评估JVET 360Lib支持的7种格式。将每个序列从其源等距柱状投影转换为编码投影,在多个码率点进行压缩、重建并转换回原始格式。采用PSNR、球形PSNR、加权球形PSNR和Bjøntegaard delta rate评估性能。还将结合投影转换、神经压缩和逆投影的可微分流水线与360Lib进行对比。结果显示,采用scale-space flow模型时,等距柱状投影和填充等距柱状投影的压缩效率最高,而基于立方图和菱形十二面体的投影效果较差。这与传统HM-16.16编解码器的结果不同,对于后者,基于立方图的格式(尤其是等角和调整后的立方图投影)优于等距柱状格式。基于光流的神经模型受益于单面投影的空间连续性,而基于块的混合编解码器更适配多面布局。这些发现表明投影效率依赖于编解码器,为学习型360度视频压缩的投影选择提供了指导。

英文摘要

360-degree video supports immersive applications such as virtual reality, autonomous driving, and education. Because spherical content cannot be processed directly by conventional video codecs, it must first be mapped to a two-dimensional projection. Projection choice affects spatial continuity, sampling uniformity, motion estimation, and compression efficiency. This thesis investigates how projection format influences end-to-end neural compression of 360-degree video. Seven formats supported by JVET 360Lib are evaluated using the scale-space flow model, JVET test sequences, and common test conditions. Each sequence is converted from its source equirectangular projection to a coding projection, compressed at multiple rate points, reconstructed, and converted back. Performance is assessed using PSNR, spherical PSNR, weighted spherical PSNR, and Bjøntegaard delta rate. A differentiable pipeline combining projection conversion, neural compression, and inverse projection is also compared with 360Lib. Results show that equirectangular and padded equirectangular projections provide the highest compression efficiency with the scale-space flow model, while cubemap-based and rhombic dodecahedron projections are less effective. This differs from the conventional HM-16.16 codec, for which cubemap-based formats, particularly equi-angular and adjusted cubemap projections, outperform equirectangular formats. Neural models based on optical flow benefit from the spatial continuity of single-face projections, whereas block-based hybrid codecs better accommodate multi-face layouts. These findings show that projection efficiency is codec-dependent and provide guidance for selecting projections for learning-based 360-degree video compression.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑