arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向混合神经视频表示的分辨率灵活解码

Resolution-Flexible Decoding for Hybrid Neural Video Representations

Taiga Hayami, Masaya Takabe, Hiroshi Watanabe

arXiv 2609.23555首次发表:更新:

发表机构

Waseda University(早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对混合神经视频表示中解码器受限于固定分辨率的问题,提出分辨率灵活解码框架,采用均匀上采样与对齐策略,并引入中间监督和难度感知采样,在UVG数据集上提升重建质量。

AI 中文摘要

神经视频表示(NVR)利用神经网络参数以及混合形式中的逐帧潜在嵌入来表示视频。尽管混合NVR可以通过使用内容自适应的潜在嵌入来提高重建质量,但其潜在空间大小和解码器上采样调度与目标帧分辨率绑定。对于高分辨率视频,这种依赖性可能需要较大且不均匀的上采样因子,并可能影响潜在嵌入与解码器之间的参数分配。在本文中,我们提出了一种面向混合NVR的分辨率灵活解码器框架。该解码器由均匀的2倍上采样阶段构成,其目标特征尺寸通过从最终输出分辨率向后追踪空间分辨率获得。在每个上采样阶段之后,特征图通过必要的极小填充或裁剪与目标尺寸对齐。为支持这一渐进式解码过程,我们进一步使用中间重建监督以及基于近期逐帧损失的、重建难度感知的帧采样策略。该框架保留了混合NVR的基本表示格式,因此可应用于不同的骨干网络。在UVG数据集上的实验表明,所提方法相较于相应的NVR基线提高了重建质量。

英文摘要

Neural video representations (NVRs) represent videos using neural network parameters and, in hybrid formulations, frame-wise latent embeddings. Although hybrid NVRs can improve reconstruction quality by using content-adaptive latent embeddings, their latent spatial sizes and decoder upsampling schedules are tied to the target frame resolution. For high-resolution videos, this dependency may require large and non-uniform upsampling factors and can affect the parameter allocation between the latent embeddings and the decoder. In this paper, we propose a resolution-flexible decoder framework for hybrid NVRs. The decoder is constructed from uniform \(2\times\) upsampling stages, whose target feature sizes are obtained by tracing the spatial resolution backward from the final output resolution. After each upsampling stage, the feature map is aligned with the target size by minimal padding or cropping when necessary. To support this progressive decoding process, we further use intermediate reconstruction supervision and a reconstruction-difficulty-aware frame sampling strategy based on recent frame-wise losses. The framework preserves the basic representation format of hybrid NVRs and can therefore be applied to different backbones. Experiments on the UVG dataset show that the proposed approach improves reconstruction quality over the corresponding NVR baselines.

CommentsAccepted to VCIP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑