发表机构
University of Ghana(加纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究广义少样本3D点云分割问题,提出无训练方法,用冻结的3D视觉语言模型与概念分割器,通过跨视图一致性协调,在ScanNet200和ScanNet++基准上提升新类mIoU,且无需少样本支持,效果优于现有方法。
AI 中文摘要
广义少样本3D点云分割(GFS-PCS)要求模型将场景分割为训练时见过的许多基础类和一组新类。当前的先进方法通过将密集但有噪声的3D视觉语言先验与少样本支持进行协调来处理新类,但这需要基础3D标签、逐集训练以及支持注释。本文提出一种无训练的方法,将冻结的3D视觉语言模型(RegionPLC)作为密集先验,与冻结的可提示概念分割器(SAM3)配对,由新类名称提示并从RGB视图中提取,通过跨视图一致性协调两者。在ScanNet200 GFS-PCS基准上,该无训练开放词汇管道将新类的mIoU提高了2.6,同时保持基础准确率在0.5以内,恢复了与使用更多监督的训练后先进状态之间三分之一的新类差距。在更难的ScanNet++基准上,相同管道使新类mIoU几乎翻倍(从16.2提高到31.9),同时基础成本仅增加1.7,将调和均值从21.5提高到31.1。此外,将少样本支持注入管道不仅没有提升效果,反而通过分类器降低了性能,因此该方法无需少样本支持。
英文摘要
Generalized few-shot 3D point-cloud segmentation (GFS-PCS) asks a model to segment a scene into many base classes seen at training time and a set of novel classes. The state of the art reaches novel classes by reconciling a dense but noisy 3D vision-language prior with the few-shot support, but it pays for this with base 3D labels, per-episode training, and the support annotations themselves. We ask how far the same reconciliation can go with none of these: no training, no 3D labels, and not even the few-shot support. We pair a frozen 3D vision-language model (RegionPLC) as a dense prior with a frozen promptable concept segmenter (SAM3), prompted by the bare novel class names and lifted from posed RGB views, and reconcile the two by cross-view consistency: a point becomes novel only when enough of the views that see it agree. On the ScanNet200 GFS-PCS benchmark this fully training-free, open-vocabulary pipeline improves novel mIoU by +2.6 over the training-free dense prior while holding base accuracy within 0.5, and recovers a third (33%) of the novel-class gap to the trained state of the art that uses far more supervision. We further show that injecting the few-shot support into the pipeline, as a fusion gate and as a prototypical dense classifier, adds nothing over consistency alone and in fact degrades it through the classifier, which is why the method needs no support at all. On the harder ScanNet++ benchmark, where the dense prior is far weaker on novel classes, the same pipeline nearly doubles novel mIoU (+15.7, from 16.2 to 31.9) at a 1.7 base cost, lifting the harmonic mean from 21.5 to 31.1