CoverPrune:基于最优传输的3D视觉语言模型的覆盖驱动型令牌剪枝
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
浏览论文内容
中文总结 AI 辅助
针对3D VLMs因大量视觉令牌导致的计算瓶颈,提出CoverPrune框架,将剪枝转化为最优传输问题,结合FST成本与SGS算法,实现高效剪枝且性能优异。
中文摘要 AI 辅助
尽管3D视觉语言模型(3D VLMs)已展现出卓越的空间推理能力,但它们存在大量视觉令牌,这在推理过程中造成了严重的计算瓶颈。现有的令牌剪枝方法主要依赖基于多样性的选择,即丢弃相似令牌以最大化分散度。然而,在3D环境中,这种方法经常会丢弃具有代表性的原型令牌,转而保留异常值,从而破坏了空间推理所必需的多视图一致性和几何结构。在本文中,我们提出了3D VLM令牌剪枝的范式转变:从最大化多样性转向保留视觉证据覆盖。我们引入了CoverPrune,这是一种无需训练的框架,它将推理时的令牌剪枝表述为最优传输(OT)问题。为了克服该表述中固有的难以处理的组合子集选择问题,我们设计了特征-空间-时间(FST)传输成本和目标容量,以及一种高效的空间引导贪心选择(SGS)算法来近似OT目标。此外,我们还提出了CoverPrune-Lite,这是一种利用空间结构化局部匹配的加速变体,开销极小。在多个3D视觉空间推理基准上进行的大量实验表明,我们的方法实现了最先进的令牌效率,即使在极具攻击性的剪枝预算下也能保持稳健的推理性能。访问我们的项目网站:this https URL。
英文摘要
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maximize dispersion. However, in 3D environments, this approach frequently drops representative prototype tokens in favor of outliers, breaking the multi-view consistencies and geometric structures essential for spatial reasoning. In this paper, we propose a paradigm shift for 3D VLM token pruning: from maximizing diversity to preserving visual evidence coverage. We introduce CoverPrune, a training-free framework that formulates inference-time token pruning as an Optimal Transport (OT) problem. To overcome the intractable combinatorial subset selection inherent in this formulation, we design the Feature-Spatial-Temporal (FST) transport cost and target capacity, along with an efficient Spatial-Guided Greedy Selection (SGS) algorithm to approximate the OT objective. Furthermore, we propose CoverPrune-Lite, an accelerated variant utilizing spatially structured local matching for minimal overhead. Extensive experiments across multiple 3D visual-spatial reasoning benchmarks demonstrate that our methods achieve state-of-the-art token efficiency, maintaining robust reasoning performance even under highly aggressive pruning budgets. Visit our project website at https://github.com/Brucess/CoverPrune.
发表机构
- Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
- LIGHTSPEED
机构由 AI 辅助整理,请以论文原文为准。