发表机构
School of Artificial Intelligence, Beijing University of Posts and Telecommunications; Shandong Computer Science Center, Qilu University of Technology; Macquarie University; Spatialtemporal AI; University of Macau; Aerospace Information Research Institute(北京邮电大学人工智能学院; 山东省计算中心(国家超级计算济南中心),齐鲁工业大学; 麦考瑞大学; 时空人工智能公司; 澳门大学; 航天信息研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言引导地理定位中现有方法不足,提出DlgPR,将场所识别转为对话驱动推理过程。构建DlgQuest-Cities基准及统一推理框架,用课程训练DQ-pilot,通过特定指标指导学习并实验,该方法显著优于基线。
AI 中文摘要
受人类交流空间信息方式的启发,语言引导的地理定位因其直观实用价值而备受关注。尽管取得了进展,但多数方法仍依赖静态一次性检索范式,难以处理现实世界自然语言描述中的模糊性和不完整性。我们提出向推理检索的范式转变,引入对话式场所识别(DlgPR),将定位视为交互式、对话驱动的推理过程。为支持此新任务,我们展示了DlgQuest-Cities,首个大规模基于对话的场所识别基准,以及一个统一推理框架,该框架将跨模态多级检索器与智能提问器DQ-pilot相结合。DQ-pilot采用课程训练:先在精心策划的DQ-cities-20k子集上进行监督微调,然后通过GRPO在更难的DQ-cities-10k分割上进行强化优化。两个任务对齐的指标指导学习:用于课程采样的判别难度指数(DDI)和直接衡量问题引起的检索改进的位置检索增益(PRG)奖励。实验表明,这种基于推理的方法显著优于基线。代码和模型可在该https网址获取。
英文摘要
Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleteness inherent in real-world natural language descriptions. We propose a paradigm shift to reasoning retrieval and introduce Dialogue Place Recognition (DlgPR), which casts localization as an interactive, dialogue-driven reasoning process. To support this new task, we present DlgQuest-Cities, the first large-scale dialogue-based benchmark for place recognition, and a unified reasoning framework that couples a cross-modal multi-level retriever with an intelligent questioner, DQ-pilot. DQ-pilot is trained in a curriculum: supervised fine-tuning on a curated DQ-cities-20k subset followed by reinforcement refinement on a harder DQ-cities-10k split via GRPO. Two task-aligned metrics guide learning: a Discriminative Difficulty Index (DDI) for curriculum sampling and a Positional Retrieval Gain (PRG) reward that directly measures retrieval improvement induced by a question. Experiments show this reasoning-based approach significantly outperforms baselines. The code and model are available at https://github.com/Graysonggg/DlgPR.
CommentsAccepted to CVPR 2026