arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoVeR:基于覆盖率的视觉令牌剪枝用于VLM中的多视图3D推理

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre

arXiv 2609.08345首次发表:更新:

发表机构

Carnegie Mellon University; Meta Reality Labs(卡内基梅隆大学; 元宇宙现实实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoVeR是一种无需训练的确定性视觉令牌剪枝方法,通过空间覆盖率选择令牌,解决多视图3D推理中冗余问题,在三个基准上超越现有方法,仅用8%令牌保留93.5%性能。

AI 中文摘要

将3D场景表示为多视图图像,使得2D VLM能够通过重用预训练的先验知识在3D中进行推理,从而规避了标注3D数据稀缺的问题。然而,这会产生数千个冗余的视觉令牌,其成本随每个视图的增加而增长。现有的视觉令牌剪枝方法分为两类,每类在多视图3D设置中均存在局限性。基于学习的重要性方法通过注意力或编码器特征对令牌进行排序;由于这里的冗余本质上是空间性的,它们会从少数显著区域保留近乎重复的令牌,而场景的大部分区域则未被表示。体素化方法改善了空间覆盖率,但无法强制执行精确的令牌预算,并且随着多视图观测在3D中重叠而饱和,将保留率限制在远低于目标水平。我们表明空间覆盖率与3D推理性能相关,并引入了CoVeR,一种确定性的、无需训练的选取器,仅使用令牌坐标,不依赖任何学习信号。CoVeR选择能够共同覆盖场景每个区域的令牌,并解决了上述两类方法的局限性:它强制执行每个场景的精确预算,打破了体素化的饱和平台,并避免了基于学习重要性方法产生的近乎重复的选择。大量实验表明,CoVeR在所有三个3D推理基准上均优于先前的最先进方法,并作为即插即用模块在四种VLM上进行了测试,展现出良好的泛化能力。值得注意的是,仅使用约8%的视觉令牌,它就能保留全令牌性能的93.5%,在基准测试中平均超过最先进方法3.9个百分点。

英文摘要

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only $\approx$8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.

Comments21 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑