arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoANeRV:坐标感知的令牌空间神经视频表示

CoANeRV: Coordinate-Aware Token-Space Neural Video Representation

Jialong Guo, Ke Liu, Mengxuan Li, Jiajun Bu, Haishuai Wang

arXiv 2608.13938首次发表:更新:

发表机构

College of Computer Science, Zhejiang University; Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems(浙江大学计算机科学学院; 浙江省无障碍感知与智能系统重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CoANeRV框架,通过轴自适应位置编码等实现高效摊销式视频表示,提升重建质量并降低内存消耗,相关代码已公开。

AI 中文摘要

视频的神经表示(NeRV)通过在网络权重中存储视频特定信息,展现出强大的重建保真度。然而,现有方案通常需要昂贵的逐视频优化或视频特定权重生成,难以扩展为高效的摊销式视频表示。我们提出CoANeRV,这是一种坐标感知的令牌空间框架,将更广泛的令牌条件神经场范式适配于摊销式视频表示。CoANeRV在一次前向传播中形成紧凑的视频令牌,并使用共享的坐标条件解码器重建连续的时空查询,避免了逐视频解码器的优化或生成,同时保留了坐标级别的重建灵活性。为使令牌空间重建有效,CoANeRV引入了一种坐标感知解码架构,通过轴自适应位置编码和温度调节的交叉注意力,将时空查询与视频令牌对齐。分块坐标查询进一步降低了峰值注意力内存,使高分辨率重建成为可能。在不同视频数据集上的实验表明,CoANeRV相比先前的前向NeRV和INR基线,始终提升了重建质量;相比基于注意力的坐标解码器,降低了峰值内存;且提供了无需逐视频优化的高效摊销式编码。这些结果支持所提出的前向令牌形成、时空坐标检索和内存受限的密集查询的视频特定组合。代码可在该https URL获取。

英文摘要

Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose CoANeRV, a coordinate-aware token-space framework that adapts the broader token-conditioned neural-field paradigm to amortized video representation. CoANeRV forms compact video tokens in one feed-forward pass and uses a shared coordinate-conditioned decoder to reconstruct continuous spatio-temporal queries, avoiding per-video decoder optimization or generation while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results support the proposed video-specific combination of feed-forward token formation, spatio-temporal coordinate retrieval, and memory-bounded dense querying. The code is available at https://github.com/jialong2023/CoANeRV.

Comments27 pages, 13 figures. Code is available at https://github.com/jialong2023/CoANeRV

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑