arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiResNets:使用中央凹残差流的原生全高清视频识别

HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams

Shivani Mall, Swarnim Jain, Joao F. Henriques

arXiv 2608.02140首次发表:更新:

发表机构

University of Cambridge; University of Oxford; VGG(剑桥大学; 牛津大学; 视觉几何组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出HiResNets,利用中央凹残差流消除残差流分辨率的二次依赖,实现快速处理高分辨率视频,在以自我为中心的视频识别任务中性能更优。

AI 中文摘要

近期图像与视频识别领域的诸多进展都以内存消耗为代价:更大的模型、更高的分辨率、更长的时间上下文。其中不可避免的一点是,基于图像分辨率的内存与计算量会呈二次(或更高)增长,这是卷积网络与视觉Transformer所采用的网格采样特性导致的。本研究中,我们探索了其卷积块内存与计算量呈对数平方增长而非二次增长的残差网络,使其能快速处理超高分辨率视频。核心思路是将残差网络的残差流作为高分辨率缓冲区,卷积块仅通过对数极坐标图像扭曲操作对其进行读写。各层会自适应地聚焦于帧的不同部分,仅在聚焦点附近保留极高分辨率。残差流中会构建出完整的高分辨率表示,类似生物视觉中眼睛扫视形成完整图像的过程,且本文提出了一种理论构造,消除了残差流分辨率的二次依赖关系。实验表明,我们提出的HiResNets能像人类视觉一样对场景进行中央凹聚焦,在困难的以自我为中心的视频识别任务中表现更优,尤其是针对包含小物体的以自我为中心视频与细粒度识别任务。

英文摘要

Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal contexts. An inevitable component is the quadratic (or larger) growth of memory and compute based on image resolution, which is a property of the grid sampling used in convolutional networks and vision transformers. In this work we study residual networks whose convolutional blocks have logarithmic-square growth instead, enabling them to process very high-resolution video quickly. The key insight is to use a residual architecture's residual stream as a high-resolution buffer, to which convolutional blocks only read and write via log-polar image warp operations. Layers adaptively focus on different parts of each frame, with very high resolution only near the focus point. A complete high-resolution representation is built up in the residual stream, analogous to eye saccades creating a complete picture in biological vision, and a theoretical construction is presented that eliminates the quadratic dependency of the residual stream resolution. Experiments demonstrate that our proposed HiResNets learn to foveate around scenes similarly to human vision, and have superior performance in difficult egocentric video recognition tasks, especially egocentric video with small objects and fine-grained recognition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑