检索到定位:弥合大语言模型与LiDAR几何用于空间定位
Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding
浏览论文内容
中文总结 AI 辅助
针对空间定位中复杂关系理解不足的问题,提出SpatialLiDAR-LM,结合LLM语言先验与LiDAR几何,通过检索和局部细化实现精确坐标预测,显著优于现有模型。
中文摘要 AI 辅助
LiDAR为空间感知任务(如自动驾驶和户外机器人中的目标检测)提供了精确的几何信息。然而,识别和定位单个目标并不足以回答需要组合空间关系并定位预期目标的问题。受自动驾驶领域大语言模型(LLMs)最新进展的启发,我们利用其语言先验来解释复杂的空间问题,并在LiDAR几何中定位所指目标。为支持这种空间定位能力,我们引入了SpatialLiDAR-QA,它将单步和多步关系定位与互补的空间理解任务相结合。我们进一步提出了SpatialLiDAR-LM,它将LiDAR点特征与LLM对齐,并通过基于语言条件、位置感知的提议检索和局部点细化来定位目标坐标。这种设计直接从局部LiDAR几何中推导目标坐标,而非通过文本语言解码。实验表明,在精确坐标预测任务上,我们相较于具有代表性的LiDAR-语言模型和多摄像头VLM取得了显著改进。我们的数据集和模型训练代码将公开发布。
英文摘要
LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.
发表机构
- Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。