arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GISAgentBench:面向GIS任务的、由从业者提供的用于评估大语言模型智能体的基准测试

GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks

Abhinav Pothuri, Zhe Jiang, Zelin Xu, Di Yang

arXiv 2608.01645首次发表:更新:

AI 中文总结

针对现有GIS智能体基准缺乏真实输出等问题,推出由从业者提供的GISAgentBench基准,含349项多步骤GIS任务,评估显示最佳智能体仅完成32.7%的任务,凸显真实GIS工作流的挑战性。

AI 中文摘要

地理信息系统(GIS)专业人员依赖多步骤空间分析工作流为城市规划、灾害响应和环境监测领域的决策提供支持。该过程繁琐、耗时且易出错。尽管近期配备外部工具的大语言模型(LLM)智能体具备自动化地理空间分析的潜力,但其执行真实GIS工作流的能力仍未得到充分探索。现有GIS智能体基准测试数据集大多源自教科书、教程或LLM生成的种子,规模和轨迹深度有限,更重要的是,没有一个数据集提供真实输出,因此依赖代码相似度、轨迹匹配或LLM与视觉语言模型(VLM)评判等替代信号,这类信号可能混淆工作流相似性与任务正确性。为解决这一缺口,我们推出GISAgentBench,该基准包含349项多步骤GIS任务,这些任务由GIS Stack Exchange精选,并在六个选定的目标地理区域的真实公共数据上实例化。每项任务均附带可执行的参考轨迹和精确的真实输出文件,支持超越LLM评判的严格、确定性、容忍度感知的输出匹配。对六个LLM模型的评估显示,真实GIS工作流仍具挑战性:在严格的容忍度感知评分下,表现最佳的智能体仅完成32.7%的任务,尽管多数模型生成的输出接近真实值。

英文摘要

Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language model (LLM) agents equipped with external tools have the potential to automate geospatial analysis, their ability to perform realistic GIS workflows remains largely unexplored. Existing GIS agent benchmarking datasets are mostly drawn from textbooks, tutorials, or LLM-generated seeds and remain limited in size and trajectory depth. More importantly, none provides ground truth outputs. They therefore rely on surrogate signals such as code similarity, trajectory matching, or LLM and VLM judges, which can conflate workflow resemblance with task correctness. To address this gap, we introduce GISAgentBench, a benchmark of 349 multi-step GIS tasks curated from GIS Stack Exchange and instantiated on real public data across six selected geographic areas of interest. Each task ships with an executable reference trajectory and an exact ground truth output file, enabling strict, deterministic, tolerance-aware output matching beyond LLM judging. Evaluations of six LLM models reveal that realistic GIS workflows remain challenging: the best agent completes only 32.7% of tasks under strict tolerance-aware scoring, although most models produce outputs that are close to the ground truth.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑