基于八叉树的视频表示
Octree-based Video Representation
浏览论文内容
中文总结 AI 辅助
针对视频均匀网格表示效率低的问题,提出基于八叉树的OctVideo表示,通过递归划分时空体积并配合轻量VAE重建,在K400和DAVIS上以最少FLOPs实现高重建质量与快速编解码,并支持高效视频理解。
中文摘要 AI 辅助
视频模型通常使用均匀网格,尽管视觉复杂度在空间和时间上差异显著。我们提出了OctVideo,它用八叉树近似一个视频片段。这种层级结构递归地将时空体积划分为八个子体积,使得平滑区域保持粗糙,而细节丰富的区域获得更精细的单元。每个叶子节点存储局部RGB值和时空梯度,并辅以轻量级学习残差。在重建时,Conv1D VAE将序列化单元映射到规则潜在网格,并在解码过程中选择性地细化细节。我们的VAE在Kinetics-400(K400)上达到了36.12 dB的PSNR,每片段参数量为38.2M,计算量为189.4 GFLOPs。它还能零样本泛化到高分辨率密集标注视频分割(DAVIS)2016数据集,重建质量与评估中最佳模型相当。在两个数据集上,它所需模型FLOPs最少,且在评估模型中编码和解码速度最快。OctVideo还支持视频理解,在从头训练时仅用少量输入令牌即可达到有竞争力的识别性能。通过利用视频信号中已有的冗余并高效处理稀疏结构,OctVideo为视频提供了一种高效表示。
英文摘要
Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.