GAPrompt++:面向3D视觉模型的多粒度几何感知点云提示
GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model
浏览论文内容
中文总结 AI 辅助
针对3D点云任务适配中现有提示方法忽略几何结构的问题,提出多粒度几何感知提示方法GAPrompt++,通过点位移、关键点提示和提示传播机制,以少于2%可训练参数超越全量微调,并构建新基准。
中文摘要 AI 辅助
预训练的3D视觉模型显著推进了点云分析,但通过全量微调将其适配到下游任务在计算上昂贵且存储密集。参数高效微调(PEFT)通过降低适配成本和存储负担提供了一种有前景的替代方案。然而,现有的基于提示的方法忽略了点云的内在几何结构,从而限制了其适配能力。这一局限源于它们无法编码细粒度几何线索和粗粒度结构语义,也未能通过模型层级有效传播此类信息。为解决这些挑战,我们提出了GAPrompt++,一种多粒度几何感知提示方法,为高效的3D任务适配提供更丰富的几何指导。具体而言,我们引入了一个点位移提示器(Point Shift Prompter),在不同尺度上提取多粒度几何特征,从而在适配期间实现实例特定的几何调整。接下来,一个关键点提示器(Keypoint Prompter)自适应地生成点级提示,以突出局部几何显著性和细粒度结构细节。此外,一种提示传播机制在整个特征提取层级中注入这些多粒度几何线索,增强捕获基本几何特征的能力。大量实验表明,GAPrompt++在基于提示的PEFT方法中达到了最先进的性能,甚至在多样基准上超越了全量微调,同时仅需少于2%的可训练参数。此外,为解决现有评估数据集的饱和问题,我们构建了两个更具挑战性的基准,源自3D高斯泼溅和多视图立体重建,提供多样且逼真的点云场景,以促进未来研究。
英文摘要
Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-Efficient Fine-Tuning (PEFT) offers a promising alternative by reducing both adaptation cost and storage burden. However, existing prompting-based approaches ignore the intrinsic geometric structures of point clouds, thereby limiting their adaptation capability. This limitation stems from their inability to encode both fine-grained geometric cues and coarse-grained structural semantics, as well as failing to propagate such information effectively through the model hierarchy. To address these challenges, we propose GAPrompt++, a multi-granular geometry-aware prompting method that provides richer geometric guidance for efficient 3D task adaptation. Specifically, we introduce a Point Shift Prompter that extracts multi-granular geometric features across different scales, enabling instance-specific geometric adjustments during adaptation. Next, a Keypoint Prompter adaptively generates point-level prompts to highlight local geometric saliency and fine-grained structural details. Furthermore, a Prompt Propagation mechanism injects these multi-granular geometric cues throughout the feature extraction hierarchy, strengthening the ability to capture essential geometric characteristics. Extensive experiments show that GAPrompt++ achieves state-of-the-art performance among prompting-based PEFT methods and even surpasses full fine-tuning across diverse benchmarks, while requiring less than 2\% trainable parameters. In addition, to address the saturation of existing evaluation datasets, we construct two more challenging benchmarks derived from 3D Gaussian Splatting and Multi-View Stereo reconstruction, offering diverse and realistic point cloud scenarios to promote future research.
发表机构
- Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
- Intelligent Science and Technology Academy of CASIC(中国航天科工集团智能科技研究院)
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- Department of Automation, Tsinghua University(清华大学自动化系)
机构由 AI 辅助整理,请以论文原文为准。