发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出MESSY STREETS基准测试集,发现商业地理编码器召回率比开源系统高49个百分点,非规范地址的候选返回率差异是主因,归一化预处理可缩小两者差距。
AI 中文摘要
我们推出MESSY STREETS,这是一个针对逐字网页地址评估地理编码器的基准测试集,包含存在性验证和表面形式差异的可控测量。与基于干净或人工扰动地址的传统基准不同,MESSY STREETS包含的地址其表面形式与规范表示存在差异,且其组成部分可能缺失、重复、格式错误或不完整。该基准测试集构建自2024年12月的Web Data Commons语料库,参考位置由OpenAddresses或OpenStreetMap建立。最强的商业地理编码器在召回率上比开源系统高出多达49个百分点。这一差距主要由非规范地址上候选返回率的差异驱动;一旦返回候选结果,各系统的位置精度大致相当。仅非规范表面形式就导致召回率损失多达25个百分点。通过检查Nominatim的查询处理流程,我们发现其合取匹配会让单个未识别的 token 使原本有效的查询失效。结果表明,地理编码器的选择对于处理嘈杂地址数据的应用而言是一项重要的设计决策,且归一化和预处理可大幅缩小开源与商业地理编码器之间的差距。
英文摘要
We introduce MESSY STREETS, a benchmark for evaluating geocoders on verbatim web addresses, with existence verification and controlled measurement of surface-form divergence. Unlike conventional benchmarks based on clean or synthetically perturbed addresses, MESSY STREETS contains addresses whose surface forms diverge from canonical representations and whose components may be missing, repeated, malformed, or incomplete. The benchmark is constructed from the December 2024 Web Data Commons corpus, with reference locations established from OpenAddresses or OpenStreetMap. The strongest commercial geocoders outperform open-source systems by up to 49 percentage points in recall. This gap is driven primarily by differences in candidate return rates on non-canonical addresses; once a candidate is returned, positional accuracy is broadly comparable across systems. Non-canonical surface form alone accounts for up to 25 percentage points of recall loss. Examining Nominatim's query-processing pipeline, we show that its conjunctive matching lets a single unrecognised token zero an otherwise valid query. The results demonstrate that geocoder choice is a consequential design decision for applications processing noisy address data, and that normalisation and preprocessing could substantially narrow the gap between open-source and commercial geocoders.