基于隐式神经表示的数据驱动视频编解码器
Data-driven Video Codec with Implicit Neural Representations
查看机构详情
- Department of Electronics and Computer Engineering, Thapathali Campus Institute of Engineering, Tribhuvan University(电子与计算机工程系,塔帕塔利校区工程学院,廷布孙大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究提出用存储单个正弦表示网络权重的方式编解码视频及音频,介绍网络结构、压缩方法等,在测试视频上有一定效果,与其他编解码器比较有不足,还展示了基于浏览器的原型。
中文摘要 AI 辅助
传统编解码器将视频存储为压缩像素数据,本文将视频及其音频轨道存储为单个正弦表示网络(SIREN)的权重,该网络将时空坐标映射到RGB值和音频幅度。网络采用单独的音频和视频初始化层、共享的全连接隐藏层堆栈以及三个输出分支。通过基于响应的知识蒸馏将过拟合的教师网络压缩为较小的学生网络,然后进行16位对称权重量化和无损LZMA2(xz)编码。在测试视频上,量化后的学生网络达到视频PSNR为28.72dB,SSIM为0.75,音频PSNR为24.18dB,对数谱距离为10.69dB,管道将表示从9.05MiB缩小到2.33MiB,总体压缩率为2.61。与H.264、HEVC和MP3进行比较,报告该方法的不足之处,并描述了一个基于浏览器的原型,可通过WebRTC训练、传输和解码这些模型。
英文摘要
A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of a single sinusoidal representation network (SIREN) that maps space-time coordinates to RGB values and audio amplitudes. The network uses separate audio and video initialization layers, a stack of shared fully connected hidden layers, and three output branches: one for video and two Siamese audio branches whose disagreement is used to estimate and subtract residual noise. The overfitted teacher network is then compressed by response-based knowledge distillation into a smaller student, followed by 16-bit symmetric weight quantization and lossless LZMA2 (xz) encoding. On a 6.08 MiB test video, the quantized student reaches a video PSNR of 28.72 dB with SSIM of 0.75, and an audio PSNR of 24.18 dB with a log spectral distance of 10.69 dB, while the pipeline shrinks the representation from 9.05 MiB to 2.33 MiB, an overall compression ratio of 2.61. A bit-width sweep from 1-bit to 32-bit quantization shows that reconstruction quality saturates at 16 bits. We compare against H.264, HEVC, and MP3, report where the approach falls short of them, and describe a browser-based prototype that trains, transfers, and decodes these models over WebRTC.