arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19911cs.CV

CitySTAR:面向开放词汇城市三维定位的结构化与拓扑感知推理

CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

  • HKUST(GZ)(香港科技大学(广州))
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Shuai Zhang, Hongye Hou, Qinghe Liu, Zhuoxiao Li, Dongli Wu, Jing Ou, Yuan Liu, Wufan Zhao

AI总结:

针对城市级开放词汇三维定位,提出免训练框架CitySTAR,通过场景图与超图拓扑验证及反思性跨模态定位实现结构化约束推理,并构建增强基准CitySTAR-3D,显著提升定位性能与可解释性。

AI中文摘要:

三维定位旨在从自然语言中定位复杂场景中的目标实体,在具身感知和空间推理中发挥着基础作用。然而,现有方法大多依赖特征相似性或直接匹配,难以将自然语言意图与十亿级城市点云中隐含的语义和几何结构联系起来。我们将城市级三维定位重新表述为结构化约束推理,其中描述语义被组织为对开放词汇三维实体、属性和空间关系的可计算跨模态约束。我们提出了CitySTAR,一个免训练框架,用于推理驱动的城市三维定位。CitySTAR将原始十亿级城市点云提升为查询就绪的开放词汇三维实例场景图,由CodeLLM驱动的工具为节点属性和三维空间关系提供多模态证据。然后,它使用配对超图对目标-上下文拓扑进行建模,并执行双向拓扑验证以进行结构消歧。最后,一个反思性跨模态定位模块整合拓扑一致性和以候选为中心的三维视觉证据,在度量感知的三维上下文图上做出决策。为进一步支持这一设置,我们引入了CitySTAR-3D,一个增强基准,提高了城市级三维定位中的语义覆盖、实例完整性、边界框保真度和空间关系复杂性。大量实验表明,CitySTAR持续改进开放世界城市三维定位,同时保持强大的可解释性和泛化能力。

英文摘要:

3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.

↑