AI 中文总结
研究人员构建了覆盖1981-2025年统计学与数据科学领域的大规模引文网络数据集StatCite,含18.9万余篇论文及四类引文网络,可用于科学文献的统计分析等研究。
AI 中文摘要
本文介绍StatCite,这是一个覆盖1981年至2025年统计学与数据科学领域出版物的大规模引文网络数据集。该数据集包含从62种代表性期刊中收集的189101篇研究论文,提供的文献元数据包括标题、作者列表、出版商、发表年份、摘要、关键词以及参考文献列表。基于收集到的出版物,我们构建了四个互补的基于引文的网络,即论文引文网络、共引网络、文献耦合网络和期刊引文网络。为说明该数据集的实用性,我们对构建的网络进行了描述性分析,并研究了论文引文网络的社区结构。结果表明,StatCite保留了大规模引文网络中常见的关键结构特征,且捕捉到统计学与数据科学领域的几个主要研究方向。通过将多种网络表示形式与丰富的文本元数据相结合,StatCite为科学文献的统计分析、知识发现和数据驱动研究提供了宝贵资源。
英文摘要
In this paper, we introduce StatCite, a large-scale citation network dataset covering publications in statistics and data science from 1981 to 2025. The dataset contains 189,101 research articles collected from 62 representative journals and provides bibliographic metadata, including title, author list, publisher, published year, abstract, keywords, and reference list. Based on the collected publications, we construct four complementary citation-based networks, namely the paper citation network, the co-citation network, the bibliographic coupling network, and the journal citation network. To illustrate the utility of the dataset, we present descriptive analyses of the constructed networks and investigate the community structure of the paper citation network. The results show that StatCite preserves key structural characteristics commonly observed in large-scale citation networks and captures several major research areas in statistics and data science. By integrating multiple network representations with rich textual metadata, StatCite provides a valuable resource for statistical analysis, knowledge discovery, and data-driven studies of scientific literature.