arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22526cs.CV

RS$^3$-Prune:用于视频目标分割的读稀疏、存储稀疏的令牌剪枝方法

RS$^3$-Prune: Read-Sparse, Store-Sparse Token Pruning for Video Object Segmentation

  • Attentive AI(注意力人工智能公司)
  • Indian Institute of Technology Delhi(印度德里理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Avilasha Mandal, Sarvesh Shashikumar

AI总结:

该研究提出无需训练的RS$^3$-Prune令牌剪枝方法,通过限制VOS流水线中跨帧注意力查询及内存库令牌范围,实现最高38.8% FPS加速、13.1%峰值内存降低,且保持竞争力的J&F指标。

AI中文摘要:

我们提出RS$^3$-Prune,这是一种无需训练的令牌剪枝方案,作为一组推理时钩子附加在现有视频目标分割(VOS)网络之上。现代VOS模型已形成一种常见的高开销设计:图像编码器为每一帧生成密集令牌网格,内存库则在所有已处理帧中累积这些令牌,以用于后续预测的条件输入。随着视频时长增加,令牌预算同时决定了每帧延迟和峰值GPU内存。因此,这些模型在长视频或在内存受限加速器上实时部署等用例中会失效。本研究认为,针对内存库型VOS的压缩应围绕令牌预算展开。RS$^3$-Prune在任意内存库型VOS流水线的两个精确位置运行:一是图像编码器与内存注意力读出器之间的边界,此处我们将参与跨帧注意力的查询限制为仅一小部分经几何信息筛选的子集;二是内存编码器与内存库之间的边界,此处我们将允许进入内存库的令牌限制为位于目标空间范围内的令牌。在多个成熟基准上,RS$^3$-Prune实现了最高38.8%的FPS加速,峰值内存使用量降低13.1%,同时与未修改的VOS网络相比,保持了具有竞争力的J&F指标。

英文摘要:

We introduce RS$^3$-Prune, a training-free token-pruning recipe that instantiates as a small set of inference time hooks atop existing video object segmentation (VOS) networks. Modern VOS models have converged on a common, expensive design: an image encoder produces a dense token grid for every frame, and a memory bank accumulates these tokens across all previously processed frames to condition future predictions. As a video grows longer, the resulting token budget governs both per-frame latency and peak GPU memory. Hence these models break on use cases such as --- long-form video or real-time deployment on memory-bounded accelerators. In this work we argue that the right axis along which to compress memory-bank VOS is the token budget itself. RS$^3$-Prune operates in two precise locations within an arbitrary memory-bank VOS pipeline: at the boundary between the image encoder and the memory-attention readout, where we restrict the queries that participate in the cross-frame attention to only a small, geometrically informed subset; and at the boundary between the memory encoder and the memory bank, where we restrict which tokens are ever permitted to enter the bank to those that lie within the object's spatial extent. Over various established benchmarks, RS$^3$-Prune delivers up to $38.8\%$ FPS speedup and reduces $13.1\%$ peak memory usage, while preserving a competitive $\mathcal{J}$&$\mathcal{F}$ compared to the unmodified VOS networks.

补充信息

↑