arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

码本容量决定分层离散视频压缩中不同分辨率下的感知质量

Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression

Manikanta Kotthapalli, Banafsheh Rekabdar

arXiv 2607.23366首次发表:更新:

发表机构

Portland State University(波特兰州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究分层离散视频压缩中感知质量的影响因素,通过实验表明感知质量主要取决于码本容量而非空间分辨率,拟合对数线性模型量化二者关系,模型在LPIPS上优于H.264和H.265,为多分辨率部署及相关模型设计提供参考。

AI 中文摘要

基于连续潜在表示的深度学习视频编解码器在调整到新空间分辨率时,通常需要针对特定分辨率重新训练或进行率失真(RD)重新校准,因为熵模型和拉格朗日权重与操作点紧密耦合。本文研究分层离散潜在编解码器是否具有相同的敏感性。通过对UCF101数据集上码本大小\(K\in\{128,256,512,1024\}\)和分辨率\(64\times64\)、\(128\times128\)、\(256\times256\)的MS - VQ - VAE视频压缩进行控制实证研究,发现感知质量(LPIPS)强烈依赖于码本容量,而对空间分辨率的依赖可忽略不计。拟合对数线性模型\(Q(K,r)=\alpha\log_2K + \beta\log_2r + \gamma\),结果表明码本容量每增加一个对数单位,其影响力约是空间分辨率的10倍。同时,底层熵效率\(\eta = H(z)/\log_2K\)随分辨率保持稳定或提高。在所有分辨率和码本大小下,模型在匹配或更低比特率时,在LPIPS上优于H.264,在\(128\times128\)时比H.264有25 - 52%的增益,在\(256\times256\)时比H.265有21 - 37%的增益。这些发现表明,在分层离散视频编解码器中,码本大小\(K\)而非空间分辨率是决定感知压缩质量的主要设计变量,这一特性可能简化多分辨率部署并为生成视频模型的可扩展离散分词器设计提供参考。

英文摘要

Learned video codecs based on continuous latent representations typically require resolution-specific retraining or rate-distortion (RD) recalibration when scaling to new spatial resolutions, because entropy models and Lagrangian weights are tightly coupled to the operating point. We investigate whether hierarchical discrete latent codecs exhibit the same sensitivity. Using a controlled empirical study of MS-VQ-VAE video compression across codebook sizes $K \in \{128,256,512,1024\}$ and resolutions $64\times64$, $128\times128$, and $256\times256$ on UCF101, we show that perceptual quality (LPIPS) depends strongly on codebook capacity but only negligibly on spatial resolution. Fitting a log-linear model $Q(K,r) = α\log_2 K + β\log_2 r + γ$ to all 12 operating points yields $α=-0.0094$ ($t=-6.6$, $p<0.001$) and $β=-0.0009$ ($t=-0.43$, $p=0.68$, not significant), with $R^2=0.82$. Codebook capacity is therefore roughly $10\times$ more influential than spatial resolution per log-unit increase. In parallel, bottom-level entropy efficiency $η=H(z)/\log_2 K$ remains stable or improves with resolution (84-87% at $64\times64$; 92-94% at $256\times256$), confirming that larger spatial grids are utilized more efficiently rather than less. Across all resolutions and codebook sizes, our models outperform H.264 on LPIPS at matched or lower bitrate, with gains of 25-52% at $128\times128$ and 21-37% over H.265 at $256\times256$. These findings suggest that codebook size $K$, not spatial resolution, is the dominant design variable governing perceptual compression quality in hierarchical discrete video codecs -- a property that may simplify multi-resolution deployment and inform the design of scalable discrete tokenizers for generative video models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑