arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自监督点Transformer的涌现式3D实例分割

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

Ted Lentsch, Santiago Montiel-Marín, Holger Caesar, Julian F. P. Kooij

arXiv 2608.15796首次发表:更新:

发表机构

Delft University of Technology; University of Alcalá(代尔夫特理工大学; 阿尔卡拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究自监督点Transformer是否含分离3D实例的结构信息,据此提出无需训练的TokenGraph3D方法,在无先验条件下大幅优于基线,使涌现的3D实例结构可见。

AI 中文摘要

室外激光雷达扫描的无监督3D实例分割传统上依赖于手工设计的几何先验,如基于密度的聚类、运动线索或投影的2D检测。本研究探究冻结的自监督点Transformer是否已包含分离对象实例所需的结构信息,而无需任何手工几何先验。将该Transformer仅用作特征提取器,在SemanticKITTI、nuScenes和Waymo Perception数据集上探究其内部表示。分析得出四个核心见解:(1)实例信号集中在注意力查询和键中,而非值或最终输出特征中;(2)输出特征在语义上坍缩,合并了查询和键保持不同的相邻同类对象;(3)该实例信号在深度上呈双峰分布,在编码器的最浅和最深阶段最强;(4)该信号主要由旋转位置编码(RoPE)驱动,移除它会消除其优势。我们将这些发现应用于方法TokenGraph3D,这是一种无需训练的分段器,通过键相似度图上的连通分量对点进行分组,既不使用基于密度的聚类也不使用邻近先验。在相同的无先验条件下,我们的方法大幅优于输出特征基线,使涌现的3D实例结构变得可见。

英文摘要

Unsupervised 3D instance segmentation of outdoor LiDAR scans has traditionally relied on handcrafted geometric priors such as density-based clustering, motion cues, or projected 2D detections. In this work, we investigate whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior. Using this transformer purely as a feature extractor, we probe its internal representations across the SemanticKITTI, nuScenes, and Waymo Perception datasets. Our analysis yields four core insights: (1) the instance signal concentrates in the attention queries and keys rather than in the values or final output features; (2) output features semantically collapse, merging adjacent same-class objects that the queries and keys keep distinct; (3) this instance signal is bimodal in depth, strongest at the shallowest and deepest encoder stages; and (4) this signal is driven predominantly by the rotary position encoding (RoPE), whose removal collapses its advantage. We put these findings into our method TokenGraph3D, a training-free segmenter that groups points via connected components on a key-similarity graph, using neither density-based clustering nor proximity priors. Under identical prior-free conditions, we substantially outperform output-feature baselines, making the emergent 3D instance structure visible.

CommentsECCV 2026 DriveX

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑