基于网络的空间上下文检索用于开放权重大语言模型:一个面向扎根地理推理的忠实性基准
Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning
浏览论文内容
中文总结 AI 辅助
针对开放权重大语言模型在地理推理中忽略外部空间上下文的问题,提出基于步行街道网络的空间上下文检索流程,并构建忠实性基准,评估模型对简报的依赖与对植入错误前提的抵抗能力。
中文摘要 AI 辅助
大型语言模型(LLMs)编码了大量的潜在地理知识,但它们在空间推理方面表现不佳,并且当仅从坐标查询时不可靠。只有当提示中提供结构化的空间上下文时,有用的行为才会出现。这提出了一个地理评估尚未考察的问题:一旦提供了正确的上下文,模型是从中推理,还是用其自身的参数记忆覆盖它?我们通过一个用于基于网络的空间上下文检索的开放流程来研究这个问题。在该流程中,所选点的周边由步行街道网络定义,即实际可步行的区域。仅使用开放数据和开放权重模型,该流程从OpenStreetMap和GHS-POP人口网格中检索特征,在代码中计算网络汇水区上的指标,并将它们作为紧凑的空间简报注入。在此基础上,我们构建了一个忠实性基准。它根据来源(基于简报或来自训练知识)和正确性标记模型做出的每一个声明,并通过一个简报驳斥的植入虚假前提来探测每个案例。我们评估了三个系列(Qwen、Gemma和Llama,其中Gemma有两代)、四个规模级别以及(如果可用)思考和非思考模式的十六种开放权重模型配置,在三个对比鲜明的城市上,对每个案例进行十次种子重采样。结果表明,对植入前提的抵抗因模型系列和代际而异,比因规模而异更明显,而简报阅读能力形成了一个部分独立的维度。这些行为未被传统的世界正确性分数或单次评估所捕获。我们在以下网址发布实现、空间简报、模型输出和声明级标签作为可复现的工作流程:此HTTP URL。
英文摘要
Large language models (LLMs) encode substantial latent geographic knowledge, yet they reason poorly over space and are unreliable when queried from coordinates alone. Useful behaviour emerges only when structured spatial context is supplied in the prompt. This raises a question geographic evaluation has left unexamined: once the right context is supplied, does the model reason from it, or override it with its own parametric recall? We take up this question with an open pipeline for network-based spatial context retriev-al. In it, the surroundings of a selected point are defined by the pedestrian street network, the area actually reachable on foot. Using only open data and open-weight models, the pipeline retrieves features from OpenStreetMap and the GHS-POP population grid, computes indicators over the network catchment in code, and injects them as a compact spatial brief. On this basis we build a faithfulness benchmark. It labels every claim a model makes by its source (grounded in the brief, or drawn from training knowledge) and its correctness, and it probes each case with a planted false premise that the brief refutes. We evaluate sixteen open-weight model configurations across three families (Qwen, Gemma and Llama, with Gemma in two generations), four size classes and, where available, both thinking and non-thinking modes, on three con-trasting cities, resampling every case over ten seeds. The results show that resistance to the planted premise varies more strongly by model family and generation than by scale, while brief-reading competence forms a partly separate dimension. These behaviours are not captured by conventional world-correctness scores or single-shot evaluation. We release the implementation, spatial briefs, model outputs, and claim-level labels as a reproducible workflow at github.com/perezjoan/NSCR-LLM.
发表机构
- Urban Geo Analytics(城市地理分析)
机构由 AI 辅助整理,请以论文原文为准。