AI 中文总结
该研究扩展人类祖先本体论(HANCESTRO),纳入代表性不足人群,构建含遗传祖先与自我报告族群的群体描述框架,以提升数据注释质量与可重复性,促进基因组学数据的整合与利用。
AI 中文摘要
成功的数据发现、整合与再利用,依赖于丰富、结构良好且机器可读的元数据,以描述数据的各个方面,从样本来源到采集过程再到实验方案。使用标准化术语以协调一致的方式表达概念,是高质量数据注释的核心,可提升数据的FAIR性,促进数据整合并推动可重复性。本文介绍人类祖先本体论(Human Ancestry Ontology, HANCESTRO),其最初开发目的是通过高层级群体描述符,改进NHGRI-EBI GWAS Catalog、人类细胞图谱(Human Cell Atlas)等遗传祖先基因组资源的标准化报告,近期已扩展至纳入基因组学与遗传学研究中多样化且此前代表性不足的人群。HANCESTRO提供群体描述符框架,涵盖基于遗传信息分析的祖先,以及基于社会文化因素、未必与遗传群体一致的自我报告族群;通过实现群体相关数据的准确且可互操作的表示,推动包容性、代表性与可重复性科学的发展。
英文摘要
Successful discovery, integration and reuse of data relies on the availability of rich, well-structured and machine-readable metadata to describe every aspect of the data, from sample sources to collection processes to experimental protocols. The use of standardised terminologies to express concepts in a harmonised fashion lies at the core of high-quality data annotation, increasing the FAIRness of the data, facilitating data integration and promoting reproducibility. Here, we describe the Human Ancestry Ontology (HANCESTRO), originally developed to improve standardised reporting of genetic ancestry genomic resources such as the NHGRI-EBI GWAS Catalog and the Human Cell Atlas through high-level population descriptors, and more recently expanded to include diverse and previously under-represented populations in genomics and genetics research. HANCESTRO provides a framework for population descriptors that includes both ancestry based on the analysis of genetic information and self-reported ethnicity, which is based on social and cultural factors that don't necessarily align with genetic populations. By enabling the accurate and interoperable representation of population-related data, it promotes inclusive, representative and reproducible science.
Comments21 pages, 1 figure. To be submitted to Cell Press Patterns