发表机构
School of Mathematical and Computational Sciences, Massey University; Joint Centre for Disaster Research, Massey University; School of Computer Science and Informatics, Cardiff University(梅西大学数学与计算科学学院; 梅西大学灾害研究联合中心; 卡迪夫大学计算机科学与信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用新西兰 Allan Herbarium 的生物标本记录识别非 gazetteer 地名(NGP),通过提取反转空间关系推导位置约束,对比确定性、概率性、LLM 方法的空间推理性能,发现概率推理精度最优,传统建模在高空间精度场景仍具优势。
AI 中文摘要
自然历史机构收集的生物标本记录构成了丰富的时间地理知识来源,捕获了不同时期记录的区域景观的生物多样性信息。本研究利用新西兰 Allan Herbarium 的数字化数据,识别这些标本地点描述中当前 gazetteer 中不存在的地名,我们将其称为非 gazetteer 地名(NGP)。这些地名通常是历史的、方言的或口语化的,在采集时被用作描述标本位置的地标。随后,我们研究仅利用标本记录中有限信息对 NGP 进行地理配准的问题。为解决这一问题,我们利用相同地名在具有不同标本位置和空间关系术语的标本记录中的重复出现,提取并反转这些关系以推导 NGP 位置的约束条件。该方法在确定性、概率性和基于 LLM 的方法中实现,从而能够对它们在基于文本的空间推理方面的优势和局限性进行比较分析。在伪 NGP 基准上,概率推理达到最高准确率(中位误差 1.43 公里;1 公里处准确率 A@1 km 为 36%),而 LLM 产生具有竞争力但精度较低的估计(中位误差 1.80 公里;1 公里处准确率 A@1 km 为 31%),这表明尽管 LLM 取得了进步,但在需要高空间精度时,传统建模仍具有优势。
英文摘要
Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times. Using digitised data from the Allan Herbarium (New Zealand), this study identifies place names in these specimen locality descriptions that are absent from current gazetteers; we refer to these as non-gazetteer place names (NGPs). These place names are typically historical, vernacular, or colloquial and were used as landmarks to describe a specimen's location at the time of collection. We then investigate the problem of georeferencing the NGPs using only the limited information available in the specimen records. To resolve this, we leverage repeated occurrences of the same place name across specimen records with different specimen locations and spatial relation terms, extracting and inverting these relations to derive constraints on NGP locations. This approach is instantiated within deterministic, probabilistic, and LLM-based methods, enabling a comparative analysis of their strengths and limitations for text-based spatial inference. On a pseudo-NGP benchmark, probabilistic inference achieves the highest accuracy (median error 1.43 km; A@1 km 36%), while the LLM yields competitive but less precise estimates (median error 1.80 km; A@1 km 31%), indicating that, despite advances in LLMs, traditional modelling remains advantageous when high spatial precision is required.
CommentsAccepted for publication in the proceedings of the Conference on Spatial Information Theory (COSIT) 2026