发表机构
University of Toulouse; IRIT(图卢兹大学; 信息与电信科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GeoBenchLLM综合基准,选取12个地理相关公开数据集评估LLMs的地理空间与时间理解能力,发现推理能力与模型规模对性能影响显著,基准可公开获取。
AI 中文摘要
在地理数据语境下,现有大语言模型(LLMs)常被置于同质化场景中研究,这极大限制了对其泛化能力的认知。本文提出GeoBenchLLM,一个用于探测大语言模型地理相关任务能力的综合基准。我们精心选取了12个来自不同地理相关任务与领域的公开数据集,利用该基准对一组大语言模型的地理空间与时间理解能力进行评估。结果显示,推理能力与模型规模对整体性能有显著影响。GeoBenchLLM可在该网址公开获取。
英文摘要
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
CommentsAccepted at CIKM2026