AI 中文总结
本研究采用自下而上的方法,从牛数据集的报告类别中提取术语语义,以在不依赖元数据或标准的情况下提高数据的互操作性和可发现性,并发现比AGROVOC更细粒度的信息。
AI 中文摘要
按年龄、性别和生产类型分类的牲畜种群数据是计算和模型的重要输入,这些计算和模型有助于我们了解全球健康状况,然而这些数据分散在不同的来源中。弥合数据孤岛以提高数据的可发现性需要互操作性。提高数据可发现性和互操作性的传统方法包括对标准化元数据进行索引。然而,创建元数据耗时且资源密集,并且通常在牲畜等领域难以实现,因为这些领域缺乏满足广泛用户群体需求的标准。当元数据存在时,通常需要根据预先存在的词汇表、本体或叙词表进行标准化,这需要一种称为“交叉映射”的技术。为了克服缺乏元数据、缺乏标准以及当前解决方案资源密集的问题,本研究采用了一种自下而上的方法。通过利用数据集中真实的报告类别,提取并分析了数据中已存在的术语的组成和语义。以牛数据为试点,我们发现来自四个数据源的五个数据集中牛术语中存在的年龄、性别和生产修饰词捕捉到了AGROVOC(世界上最大的农业词汇表)中不存在的粒度。我们讨论了这些术语的组成和语义如何用于提高数据的互操作性和可发现性,而无需首先生成元数据或创建标准。这种方法不是强制数据集符合现有词汇表,而是利用数据集中已嵌入的术语语义,使系统能够使数据更易于发现和互操作,同时保留文化和数据集特定的术语。
英文摘要
Livestock population data disaggregated by age, sex, and production are important inputs to calculations and models that inform our understanding of global health, yet these data are fragmented across disparate sources. Bridging data siloes to improve the findability of data requires interoperability. Conventional approaches to improving the findability and interoperability of data include indexing standardized metadata. However, creating metadata is time and resource-intensive and is often difficult in domains such as livestock, which lack standards that address the needs of broad user groups. When metadata exist, they typically need to be standardized against a pre-existing vocabulary, ontology, or thesaurus, requiring a technique known as `crosswalking'. To overcome issues in the absence of metadata, the lack of standards, and the resource-intensive solutions that currently exist, this study uses a bottom-up approach. By leveraging real-world reporting categories in datasets, the composition and semantics of terms already present in the data were extracted and analyzed. Using cattle data as a pilot, we find the age, sex, and production modifiers present across cattle terms from five datasets from four data sources capture granularity not present in AGROVOC, the largest agricultural vocabulary in the world. We discuss how the composition and semantics of these terms can be used to improve the interoperability and findability of data without first requiring metadata to be generated or standards to be created. Rather than forcing datasets to conform to an existing vocabulary, this approach uses the semantics embedded in terms already present in datasets, allowing systems to make data more discoverable and interoperable while maintaining culturally and dataset-specific terminology.