arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27958cs.CVcs.LG

ScoutNeRV:通过ScoutNet快速编码基于网格的视频隐式神经表示

ScoutNeRV: Rapid Encoding of Grid-Based Video INRs via ScoutNet

  • Islamic Azad University, Neyshabur(伊斯兰阿扎德大学内沙布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

Naser Alizada, Farhang Baghban, Hashem Pishkar, Ali Mousavi

中文总结 AI 辅助

ScoutNeRV通过轻量级侦察网络和记忆库中的专家初始化,加速分层视频INRs的优化,在保持竞争力的同时实现9.25倍加速。

中文摘要 AI 辅助

隐式神经表示(INRs)已成为视频压缩的一种有前景的范式,提供紧凑的神经表示,并具有灵活的空间和时间重建能力。诸如HiNeRV之类的分层基于网格的架构实现了强大的率失真性能,但需要对每个视频进行大量优化,导致高昂的编码成本。为解决这一限制,我们提出了ScoutNeRV,一种内容自适应初始化框架,用于加速分层视频INRs的优化。ScoutNeRV采用一个轻量级的、离线训练的侦察网络,该网络分析少量采样帧,并通过硬路由从记忆库中选择合适的预训练专家。然后,所选专家的分层网格和解码器参数被转移,以初始化目标HiNeRV模型,然后进行特定于视频的微调。在未见过的ReadySetGo序列上,ScoutNeRV实现了初始PSNR为34.95 dB,而标准初始化为13.70 dB,对应于微调前21.25 dB的提升。仅经过37个epoch,ScoutNeRV达到36.92 dB,并在评估的率失真配置中保持在300-epoch HiNeRV基线的0.42–0.80 dB范围内。此外,所提出的初始化在报告的运行时实验中实现了9.25倍的墙钟加速。这些结果表明,内容感知的专家初始化可以显著降低分层视频INRs的优化成本,同时保持竞争力的重建和压缩性能。代码可在https://github.com/nasserdeveloper/ScoutNeRV获取。

英文摘要

Implicit neural representations (INRs) have emerged as a promising paradigm for video compression, providing compact neural representations with flexible spatial and temporal reconstruction. Hierarchical grid-based architectures such as HiNeRV achieve strong rate--distortion performance, but require extensive per-video optimization, resulting in high encoding costs. To address this limitation, we propose ScoutNeRV, a content-adaptive initialization framework for accelerating the optimization of hierarchical video INRs. ScoutNeRV employs a lightweight, offline-trained scout network that analyzes a small number of sampled frames and selects a suitable pre-trained expert from a memory bank through hard routing. The hierarchical grid and decoder parameters of the selected expert are then transferred to initialize the target HiNeRV model before video-specific fine-tuning. On the unseen ReadySetGo sequence, ScoutNeRV achieves an initial PSNR of $34.95$~dB, compared with $13.70$~dB for standard initialization, corresponding to a $21.25$~dB improvement before fine-tuning. After only 37 epochs, ScoutNeRV reaches $36.92$~dB and remains within $0.42$--$0.80$~dB of the 300-epoch HiNeRV baseline across the evaluated rate--distortion configurations. Furthermore, the proposed initialization achieves a $9.25\times$ wall-clock speedup in the reported runtime experiment. These results demonstrate that content-aware expert initialization can substantially reduce the optimization cost of hierarchical video INRs while retaining competitive reconstruction and compression performance. The code is available at https://github.com/nasserdeveloper/ScoutNeRV.

↑