arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11674cs.AI

地理空间人工智能、Dataverse 元数据与基于地方的政府研究

Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government

  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Danny EBanks, Devika Jain

中文总结 AI 辅助

本研究利用哈佛 Dataverse 的元数据构建知识图谱,组织超十万数据集,识别地理空间与政策相关性,并指出地点解析是核心障碍,为 AI 驱动的元数据增强提供基础。

中文摘要 AI 辅助

哈佛 Dataverse 托管了超过 15 万个研究数据集,但这些数据集携带的地理信息由存储者以自由文本形式输入,且从未被整合为可搜索的结构。我们从该存储库的公共数据和元数据中构建了一个知识图谱,将 102,650 个数据集组织在一个包含 215,985 个节点和 528,003 条边的网络中,这些边将数据集与关键词、出版物、学科、期刊和地点连接起来。在这些数据集中,43,991 个(占 42.9%)至少携带一个地理空间字段、地理覆盖范围、地理单元或边界框,且 96.9% 的节点位于单一连通分量中,因此即使数据集的地理空间元数据毫无共同之处,它们之间仍保持可达性。一次保守的关键词搜索识别出 7,654 个(占 17.4%)带有地理空间标签的数据集与政策直接相关,其中选举和立法机构是最大的聚类,其次是政府行政、健康政策、交通和教育。五个数据集展示了这些元数据在政策领域和空间尺度上的行为,一个扩展用例展示了社区语言模型、结合地理聚合的立场检测以及党派语言桥接工具如何将话语与地点关联起来。核心障碍是地点解析:同一地点以多个不连通的节点出现。我们认为该图谱为开发人工智能驱动的元数据增强和实体解析提供了具体场景,并记录了其覆盖范围偏向美国城市级数据的问题。

英文摘要

Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.

↑