GeoNatureAgent (GNA):面向地理空间与环境任务的使用工具智能体生产前评估框架与基准
GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
浏览论文内容
中文总结 AI 辅助
提出GeoNatureAgent框架与103任务基准,通过固定工具MCP接口评估9个LLM,发现Claude Sonnet 4能力最高,DeepSeek V3.2性价比最优,并揭示评分严格性对结果差距的影响。
中文摘要 AI 辅助
在将使用工具的LLM智能体部署到环境和地理空间工作流之前,团队需要证据表明智能体能够针对真实API可靠地选择正确的操作。我们引入了GeoNatureAgent (GNA),一个用于使用工具智能体生产前评估的框架:一个固定的十六工具地理空间接口,以模型上下文协议(MCP)服务器形式发布,因此被测智能体是唯一变量,针对相同的工具层、任务套件和确定性评分器进行评分。其旗舰实例是一个103任务基准(一个涵盖18个类别的93任务主套件加上一个十任务比较扩展),针对一个开放的、可自托管的、服务于西班牙和葡萄牙三个环境指标的地理空间API进行评估。我们在三个温度1.0的随机种子下评估了九个LLM,将能力和每次案例成本作为正交轴进行报告。(1)Claude Sonnet 4达到最高能力(在所有103个任务上为61.7% ± 0.7%;在主套件上为60.8%),紧随其后的是DeepSeek V3.2(57.9%),而其他模型均未超过53%;(2)成本-准确率帕累托前沿主要由开放权重模型占据,DeepSeek V3.2以11.3倍更低的标价成本提供了Claude能力的93%;(3)在严格的全检查评分下,最佳模型比通用GIS基准报告的85-97%低24-36个百分点,而前四名模型的每检查部分信用(86-90%)是可比的,因此这一差距很大程度上反映了评分严格性而非仅任务难度。MCP服务器、评估框架、基准和API均可公开获取;更换工具执行器和任务套件即可为任何地理空间领域实例化等效基准。
英文摘要
Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude's capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.
发表机构
- Universidad Católica de Ávila (UCAV)(阿维拉天主教大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。