地理空间元数据通过跨学科数据集连接提升可发现性
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
浏览论文内容
中文总结 AI 辅助
本研究利用哈佛Dataverse数据,通过地理空间元数据丰富化,将跨学科数据集连接比例从58.5%提升至63.2%,证明地理元数据比关键词更可靠地促进跨学科互操作性。
中文摘要 AI 辅助
研究数据存储库是科学探究以及确保数据集遵循FAIR(可查找、可访问、可互操作、可重用)原则的重要基础设施。然而,存储库的复用取决于地理空间和主题元数据的质量与完整性,而这些元数据通常由研究人员自愿提供。鉴于有限的策展资源,即使是世界上最大的通用研究存储库哈佛Dataverse,也包含许多不完整的元数据记录,这并不令人意外。缺失字段代表信息丢失并降低互操作性。我们发现,缺失元数据较多的数据集获得的下游引用较少,与其他数据集的可解析连接也较少。这对地理空间数据集尤为重要:只有0.3%的研究数据集包含边界框,且大多数表示档案点而非完整的地理形状。我们的分析表明,地理空间元数据有助于跨学科连接概念。在将哈佛Dataverse数据集嵌入元数据知识图谱后,我们发现数据集通过共享地理空间元数据跨学科连接的可能性是通过关键词连接的两倍。这表明地理元数据比关键词词汇表(通常仍局限于特定学科)更可靠地支持跨学科互操作性。我们使用哈佛Dataverse的数据集训练并微调了一个小型语言模型。通过地理空间元数据丰富化,我们将来自不同学科的数据集通过元数据元素连接的比例从58.5%提高到63.2%。
英文摘要
Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, and Reusable) principles. However, repository reuse depends on the quality and completeness of geospatial and thematic metadata, which researchers generally provide voluntarily. Given limited curation resources, it is unsurprising that even Harvard Dataverse, the world's largest general-purpose research repository, contains many incomplete metadata records. Missing fields represent lost information and reduce interoperability. We find that datasets with more missing metadata receive fewer downstream citations and have fewer resolvable connections to other datasets. The implications are particularly important for geospatial datasets: only 0.3% of research datasets include a bounding box, and most represent archival points rather than complete geographic shapes. Our analysis shows that geospatial metadata helps connect concepts across disciplines. After embedding Harvard Dataverse datasets in a metadata knowledge graph, we find that datasets are twice as likely to connect across scientific disciplines through shared geospatial metadata as through keywords. This suggests that geographic metadata is a more reliable basis for cross-disciplinary interoperability than keyword vocabularies, which often remain discipline-specific. We train and fine-tune a small language model using datasets from Harvard Dataverse. Through geospatial metadata enrichment, we increase the share of datasets from different disciplines connected through metadata elements from 58.5% to 63.2%.
发表机构
- Institute for Quantitative Social Science, Harvard University(哈佛大学定量社会科学研究所)
- Center for Geographic Analysis, Harvard University(哈佛大学地理分析中心)
机构由 AI 辅助整理,请以论文原文为准。