发表机构
Norwegian University of Science and Technology (NTNU)(挪威科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TRACKGRAPH通过图像空间跟踪实现在线开放词汇3D场景图构建,在关键帧间传播掩码,减少VL推理开销,在多个基准上达到先进性能并显著提升速度与内存效率。
AI 中文摘要
开放词汇3D地图使机器人能够使用自然语言推理先前未知的环境。然而,现有系统通常分割每一帧输入图像,将检测结果与持久化的3D片段关联,并频繁执行昂贵的视觉-语言(VL)推理。我们提出TRACKGRAPH,一种在线开放词汇系统,在将片段融合到3D之前,直接在图像流中维护短期2D掩码身份。FastSAM掩码和CLIP特征在稀疏关键帧处计算,而密集的DINOv3特征用于在关键帧之间以高频率传播掩码。由此产生的跟踪掩码被融合到分层场景图中的类无关3D片段层中,3D关联处理跟踪中断和长期重访。紧凑的多视图CLIP嵌入实现开放词汇检索。在Replica、ScanNet++和HM3D上,TRACKGRAPH在开放词汇分割和检索方面达到了与最先进建图方法相当的性能,包括在Replica上最高的同义词频率(0.50)。在相同的NVIDIA A100上,它比ViT-H OVI-MAP快1.7倍,GPU内存使用少3.3倍。真实世界的四足机器人部署展示了7.5Hz的机载场景图构建和物体搜索,而记录的无人机数据用于测试空中视角下的方法。
英文摘要
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.