HiSC:用于高效3D场景理解的分层空间聚类令牌压缩
HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding
浏览论文内容
中文总结 AI 辅助
本文提出无训练框架HiSC,通过SGraM策略和SCluP范式实现3D VLMs的分层空间聚类令牌压缩,在高剪枝率下实现超90%令牌减少且性能下降极小,验证了其有效性。
中文摘要 AI 辅助
3D视觉语言模型(3D VLMs)可对多视图场景进行空间推理,但因存在重复观测和大量无信息区域,会产生大量令牌冗余,导致计算成本过高。尽管视觉令牌压缩在加速2D VLMs方面已展现潜力,但它无法捕捉3D场景的结构化特性,会导致空间覆盖不完整并丢失细粒度细节。本文提出HiSC,一种用于3D VLMs的无训练分层空间聚类令牌压缩框架。HiSC将令牌压缩从令牌级选择提升至聚类级处理,通过结合几何与语义线索将令牌组织为空间关联的聚类。具体而言,我们首先引入基于空间图的合并(SGraM)策略,将跨视图冗余建模为空间连通性并整合物理一致的区域,在大语言模型(LLM)推理前有效合并极其相似的冗余令牌。随后,我们在LLM推理中提出基于空间聚类的剪枝(SCluP)范式,在聚类间及聚类内执行分层压缩,在保留重要区域细粒度细节的同时维持目标实例的完整性。在多种3D推理基准上开展的大量实验验证了HiSC的有效性,尤其在高视觉令牌剪枝率下表现突出;此外,HiSC实现了超过90%的令牌减少,且性能下降极小。代码可通过此URL获取。
英文摘要
3D vision-language models (3D VLMs) enable spatial reasoning over multi-view scenes but suffer from substantial token redundancy due to duplicated observations and large uninformative regions, leading to high computational cost. Although visual token compression has shown promise in accelerating 2D VLMs, it fails to capture the structured nature of 3D scenes and leads to incomplete spatial coverage and loss of fine-grained details. In this paper, we propose \textbf{HiSC}, a training-free framework for hierarchical spatial clustering token compression in 3D VLMs. HiSC lifts token compression from token-level selection to cluster-level processing by organizing tokens into spatially grounded clusters using joint geometric and semantic cues. Specifically, we first introduce a \textbf{spatial graph-based merging (SGraM) strategy} that models cross-view redundancy as spatial connectivity and consolidates physically consistent regions, effectively merging extremely similar redundant tokens prior to LLM inference. We then propose a \textbf{spatial clustering-based pruning (SCluP) paradigm} within LLM inference, which performs hierarchical compression across clusters and within clusters, preserving object instance completeness while retaining fine-grained details for important regions. Extensive experiments on diverse 3D reasoning benchmarks show validate the effectiveness of HiSC, particularly under high visual token pruning ratios. Besides, HiSC achieves over 90\% token reduction with minimal performance degradation. Code is accessible at https://github.com/elecreak/HiSC.
发表机构
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。