数据集再利用与颠覆性人工智能研究
Dataset repurposing and disruptive AI research
浏览论文内容
中文总结 AI 辅助
本研究基于重组新颖性和变革性创造力框架,分析超万篇机器学习论文,发现数据集再利用虽短期可见度低,但关联更高颠覆性,且被后续采用时提升引用影响,多见于经验丰富、机构声望高及产学研合作的团队。
中文摘要 AI 辅助
技术进步正在推动科学各领域日益系统化和大规模的数据收集,从而驱动科学创新。特别是,人工智能研究体现了这一趋势,通过汇集用于训练和评估机器学习模型的大规模数据集而快速发展。然而,对数据日益增长的需求、创建高质量数据集的困难以及人工智能研究中易获取数据源的枯竭,引发了关于如何通过重组和再利用来最大化现有数据集价值的重要问题。在此,我们借鉴两个理论框架——重组新颖性和变革性创造力——来审视数据再利用的实践及其科学影响。聚焦于人工智能,我们分析了超过10,000篇机器学习论文中与数据再利用相关的科学成果。首先,我们发现尽管大多数被再利用的数据集在短期内未获得广泛关注,但数据再利用与更高的颠覆性相关。其次,当被再利用的数据被后续研究采用时,该再利用论文与更高的颠覆性和增加的引用影响相关联。第三,进行再利用的团队往往更有经验、机构声望更高,并涉及学术界与工业界的合作。然而,团队特征难以预测哪些被再利用的数据集将被社区采用。这些发现表明,数据再利用可能是科学发现的一种重要途径,而其成功采用在更大的团队和跨越学术界与工业界的合作中更为常见。
英文摘要
Technological advancements are enabling increasingly systematic and large-scale data collection across all areas of science, driving scientific innovation. In particular, AI research exemplifies this trend, having advanced rapidly through the assembly of massive datasets used to train and evaluate machine learning models. However, the escalating demand for data, the difficulty of creating high-quality datasets, and the exhaustion of easily accessible data sources in AI research raise important questions about how to maximize the value of existing datasets through recombination and repurposing. Here, we draw on two theoretical frameworks---recombinational novelty and transformational creativity---to examine the practice of data repurposing and its scientific impact. Focusing on AI, we analyze scientific outcomes associated with data repurposing across more than 10,000 machine learning papers. First, we find that although most repurposed datasets do not achieve broad visibility in the short term, data repurposing is associated with greater disruption. Second, when repurposed data is adopted by subsequent research, the repurposing paper is associated with higher disruption and increased citation impact. Third, repurposing teams tend to be more experienced, more institutionally prestigious, and involve academic--industry collaboration. However, team characteristics poorly predict which repurposed datasets will be adopted by the community. These findings suggest that data repurposing may be an important approach to scientific discovery, and that its successful adoption is more common among larger teams and collaborations spanning academia and industry.
发表机构
- School of Data Science, University of Virginia(弗吉尼亚大学数据科学学院)
- School of Information, University of Michigan(密歇根大学信息学院)
- Center for the Study of Complex Systems, University of Michigan(密歇根大学复杂系统研究中心)
- Computer Science and Engineering Division, University of Michigan(密歇根大学计算机科学与工程系)
- College of Information Science, University of Arizona(亚利桑那大学信息学院)
机构由 AI 辅助整理,请以论文原文为准。