arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoBenchLLM:评估大语言模型地理相关任务的综合基准

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine

arXiv 2608.07411首次发表:更新:

发表机构

University of Toulouse; IRIT(图卢兹大学; 信息与电信科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出GeoBenchLLM综合基准,选取12个地理相关公开数据集评估LLMs的地理空间与时间理解能力,发现推理能力与模型规模对性能影响显著,基准可公开获取。

AI 中文摘要

在地理数据语境下,现有大语言模型(LLMs)常被置于同质化场景中研究,这极大限制了对其泛化能力的认知。本文提出GeoBenchLLM,一个用于探测大语言模型地理相关任务能力的综合基准。我们精心选取了12个来自不同地理相关任务与领域的公开数据集,利用该基准对一组大语言模型的地理空间与时间理解能力进行评估。结果显示,推理能力与模型规模对整体性能有显著影响。GeoBenchLLM可在该网址公开获取。

英文摘要

In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.

CommentsAccepted at CIKM2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑