arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

3DZip:面向3D视觉问答的空间感知特征多样性引导的Token压缩方法

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

Changwoo Baek, Kyeongbo Kong

arXiv 2608.01185首次发表:更新:

发表机构

Pusan National University(釜山国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对3D视觉问答中Token压缩忽略空间特性的问题,提出三阶段框架3DZip,在三个基准上以128个Token保留94.7%性能、提速1.92倍,优于现有方法。

AI 中文摘要

近期的3D视觉语言模型(3D VLMs)通过将2D视觉特征投影到世界坐标来构建几何感知Token,从而支持3D视觉问答等任务的空间推理。但该设计会为每个场景生成数千个Token,带来大量计算与内存开销。尽管Token压缩已在2D VLMs中被广泛研究,现有方法依赖语义相关性或基于注意力的选择,忽略了3D Token的结构化空间特性。此外,仅靠空间邻近性无法解决3D表示中的冗余问题,因为空间聚合后仍存在对象级Token不平衡。为解决该问题,本文提出3DZip,一种三阶段Token压缩框架:首先应用粗体素化去除点级冗余,再通过行列式点过程(Determinantal Point Process)基于特征空间多样性选择锚定Token,最后在空间约束下合并剩余Token以保持几何一致性。在三个3D视觉问答基准上的实验表明,3DZip在仅使用128个Token时,可保留原始性能的94.7%,推理速度提升1.92倍,且始终优于现有压缩方法。

英文摘要

Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering. However, this design generates thousands of tokens per scene, resulting in substantial computational and memory overhead. While token compression has been extensively studied in 2D VLMs, existing approaches rely on semantic relevance or attention-based selection that overlook the structured spatial nature of 3D tokens. Moreover, redundancy in 3D representations cannot be resolved by spatial proximity alone, as object-level token imbalance persists even after spatial aggregation. To address this, we propose 3DZip, a three-stage token compression framework that first applies coarse voxelization to remove point-level redundancy, then selects anchor tokens based on feature-space diversity via a Determinantal Point Process, and finally merges remaining tokens under spatial constraints to preserve geometric coherence. Experiments on three 3D question answering benchmarks demonstrate that 3DZip consistently outperforms existing compression methods, retaining 94.7% of the original performance with only 128 tokens, achieving a $1.92\times$ faster inference speed.

CommentsAccepted to ECCV 2026. Project page: https://cvsp-lab.github.io/3DZip

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑