探索数据的社会生命:寻找你可信赖的数据
Exploring the Social Life of Data: Finding Data You Can Trust
浏览论文内容
中文总结 AI 辅助
本文提出数据使用图作为科学数据基础设施新层,结合数据使用证据与多维度信息,在国家数据平台实现原型服务,以解决可信数据查找问题。
中文摘要 AI 辅助
人工智能正在改变科学探究的规模与节奏,如今模型能够对远超任何单个研究人员所熟悉的数据仓库的数据进行搜索、整合与推理。然而这种扩展带来了一个新的问题:在模型生成可信的科学结果之前,必须先找到与研究问题相适配、对预期分析而言足够可靠,且附带足够上下文以支持负责任解读的数据。随着数据日益丰富,寻找数据的挑战已转变为寻找可信赖数据的挑战。本文探讨了数据在研究中被使用时积累的社会与经验证据如何可类比于社会信任网络,用于确定数据对特定用途的适配性与可信度。具体而言,本文将数据使用图作为科学数据基础设施的新层进行探索,该图将数据集与生成和使用它们的出版物、人员、机构、主题、软件、模型、工作流及其他数据集相连接。这些连接揭示了数据的“社会生命”:谁依赖了某个数据源、针对哪些问题、以何种组合、使用何种方法以及产生了何种可观测影响。它们可将分散的实践痕迹转化为数据使用描述符,补充传统元数据并支持对可信度与用途适配性的判断。核心主张并非是流行度确立了可信度,而是该可信度可通过适当的上下文历史来建立。因此,使用证据必须与生产质量、来源、治理、语义清晰度及社区验证相结合。本文通过在国家数据平台(NDP)内实现原型数据洞察发现服务,证明了数据使用图的可行性与价值。
英文摘要
Artificial intelligence is changing the scale and tempo of scientific inquiry. Models can now search, integrate, and reason over data far beyond data repositories familiar to any individual researcher. Yet this expansion creates a prior problem: before a model can produce a trustworthy scientific result, it must locate data that are appropriate for the question, sufficiently reliable for the intended analysis, and accompanied by enough context to support responsible interpretation. As data becomes increasingly abundant, the challenge of finding data has been overcome by the challenge of finding data that you can trust. This paper explores how the social and empirical evidence that accumulates when data are used in research can be used, analogous to social trust networks, to determine fit for purpose and trust. Specifically, the paper explores data-usage graphs as a new layer of scientific data infrastructure. A data-usage graph connects datasets to the publications, people, institutions, topics, software, models, workflows, and other datasets through which they are produced and used. These connections reveal the {\it social life of data:} who has relied on a source, for which questions, in what combinations, with which methods, and with what observable impact. They can turn scattered traces of practice into data-usage descriptors that complement conventional metadata and support judgments of trust and fitness for purpose. The central claim is not that popularity establishes trust, but that this can be grown with appropriate contextual history. Usage evidence must therefore be combined with production quality, provenance, governance, semantic clarity, and community validation. The feasibility and value of data usage graphs is demonstrated by implementing the prototype data insights discovery service within the National Data Platform (NDP).