发表机构
University of Michigan; Addis Ababa University; San Jose State University; Skyline High School; University of Pennsylvania; Carnegie Mellon University(密歇根大学; 亚的斯亚贝巴大学; 圣何塞州立大学; 斯凯莱恩高中; 宾夕法尼亚大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出AtlasNLP国家感知图集,含13000余条NLP数据集记录及两个子数据集,发现各国各任务数据集覆盖不均等问题,推动NLP评估补充地理元数据。
AI 中文摘要
理解NLP数据集中包含哪些国家的信息,对于识别数据缺口、定向数据收集、衡量进展以及为AI政策提供依据至关重要。然而,地理元数据极少可用,国家层面的表示往往隐藏在宽泛的语言层面表述之后。我们推出AtlasNLP,这是一个包含超过13000条NLP数据集记录的国家感知图集,涵盖标准化的NLP任务类别,同时跟踪所表示的人群以及数据集的生成地点。AtlasNLP包含AtlasNLP-Gold(人工整理的参考集)和AtlasNLP-Core(源自ACL的大规模数据集集合)。利用该资源,我们表明:(1)各国和各任务的数据集覆盖情况极不均衡;(2)数据集的生成与表示存在地理不对称性;(3)语言覆盖并不意味着地理覆盖。这些发现揭示了当前数据集文档实践中的盲点,并推动为国家感知的NLP评估提供更明确的地理元数据。
英文摘要
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.
CommentsProceedings of the 2026 Conference on Empirical Methods in Natural Language Processing