arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DialogueVPR:迈向对话式视觉场所识别

DialogueVPR: Towards Conversational Visual Place Recognition

Yukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu, Wenhao Xu, Shunpeng Chen, Yu Zhang, Ke Zhang, Rongtao Xu, Xuxiang Feng, Pengyang Wang

arXiv 2607.14115首次发表:更新:

发表机构

School of Artificial Intelligence, Beijing University of Posts and Telecommunications; Shandong Computer Science Center, Qilu University of Technology; Macquarie University; Spatialtemporal AI; University of Macau; Aerospace Information Research Institute(北京邮电大学人工智能学院; 山东省计算中心(国家超级计算济南中心),齐鲁工业大学; 麦考瑞大学; 时空人工智能公司; 澳门大学; 航天信息研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对语言引导地理定位中现有方法不足,提出DlgPR,将场所识别转为对话驱动推理过程。构建DlgQuest-Cities基准及统一推理框架,用课程训练DQ-pilot,通过特定指标指导学习并实验,该方法显著优于基线。

AI 中文摘要

受人类交流空间信息方式的启发,语言引导的地理定位因其直观实用价值而备受关注。尽管取得了进展,但多数方法仍依赖静态一次性检索范式,难以处理现实世界自然语言描述中的模糊性和不完整性。我们提出向推理检索的范式转变,引入对话式场所识别(DlgPR),将定位视为交互式、对话驱动的推理过程。为支持此新任务,我们展示了DlgQuest-Cities,首个大规模基于对话的场所识别基准,以及一个统一推理框架,该框架将跨模态多级检索器与智能提问器DQ-pilot相结合。DQ-pilot采用课程训练:先在精心策划的DQ-cities-20k子集上进行监督微调,然后通过GRPO在更难的DQ-cities-10k分割上进行强化优化。两个任务对齐的指标指导学习:用于课程采样的判别难度指数(DDI)和直接衡量问题引起的检索改进的位置检索增益(PRG)奖励。实验表明,这种基于推理的方法显著优于基线。代码和模型可在该https网址获取。

英文摘要

Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleteness inherent in real-world natural language descriptions. We propose a paradigm shift to reasoning retrieval and introduce Dialogue Place Recognition (DlgPR), which casts localization as an interactive, dialogue-driven reasoning process. To support this new task, we present DlgQuest-Cities, the first large-scale dialogue-based benchmark for place recognition, and a unified reasoning framework that couples a cross-modal multi-level retriever with an intelligent questioner, DQ-pilot. DQ-pilot is trained in a curriculum: supervised fine-tuning on a curated DQ-cities-20k subset followed by reinforcement refinement on a harder DQ-cities-10k split via GRPO. Two task-aligned metrics guide learning: a Discriminative Difficulty Index (DDI) for curriculum sampling and a Positional Retrieval Gain (PRG) reward that directly measures retrieval improvement induced by a question. Experiments show this reasoning-based approach significantly outperforms baselines. The code and model are available at https://github.com/Graysonggg/DlgPR.

CommentsAccepted to CVPR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑