面向3D场景理解的分区不变调优
Partition-Invariant Tuning for 3D Scene Understanding
- ZJU(浙江大学)
- GUT(桂林理工大学)
- THU(清华大学)
- PKU(北京大学)
- XJTU(西安交通大学)
- GUANGMING LAB(光明实验室)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对场景级点云理解,提出分区不变调优框架PointPiT,结合场景感知结构适配器和梯度子空间优化,以少于1%参数达到全量微调性能。
中文摘要 AI 辅助
场景级点云理解由于多样的几何形状和空间布局仍然具有挑战性。虽然预训练的3D点云基础模型(PFMs)提供了强大的可迁移性,但全量微调(FFT)会产生大量的计算和存储成本。参数高效微调(PEFT)提供了一种有前景的替代方案,但现有的PEFT方法主要关注对象级点云,忽视了大规模场景中序列化引起的分区变化。为了解决这个问题,我们提出了PointPiT,一种面向场景级点云的分区不变调优框架。具体来说,场景感知结构适配器(SSA)将局部几何模式与全局场景上下文相结合,以减轻分区引起的表示偏移。此外,梯度子空间优化(GSO)选择信息丰富且分区稳定的更新方向,在优化过程中抑制分区相关的变化。在多个场景级基准上的大量实验表明,PointPiT以不到骨干网络参数的1%实现了与全量微调相当甚至更优的性能,同时在代表性PEFT方法中取得了持续的最先进性能。
英文摘要
Scene-level point cloud understanding remains challenging due to diverse geometries and spatial layouts. While pre-trained 3D point cloud foundation models (PFMs) offer strong transferability, full fine-tuning (FFT) incurs substantial computational and storage costs. Parameter-efficient fine-tuning (PEFT) provides a promising alternative, but existing PEFT methods largely focus on object-level point clouds and overlook serialization-induced partition variations in large-scale scenes. To address this issue, we propose PointPiT, a partition-invariant tuning framework for scene-level point clouds. Specifically, a Scene-aware Structural Adapter (SSA) integrates local geometric patterns with global scene context to mitigate partition-induced representation shifts. Moreover, Gradient Subspace Optimization (GSO) selects informative and partition-stable update directions, suppressing partition-dependent variations during optimization. Extensive experiments across multiple scene-level benchmarks demonstrate that PointPiT achieves competitive or even superior performance to full fine-tuning with less than 1% of backbone's parameters, while achieving consistent state-of-the-art performance among representative PEFT methods.