学术出版网络的图方法:基于OpenAlex开放数据的异质模型与结构筛选
A Graph Approach to the Academic Publishing Network: A Heterogeneous Model and Structural Screening over OpenAlex Open Data
浏览论文内容
中文总结 AI 辅助
该研究基于OpenAlex开放数据构建异质多变量图模型,开发异常出版模式筛选方法,在两个语料库上验证其有效性,发布开源工具,可用于学术出版不端行为检测。
中文摘要 AI 辅助
学术出版生态是由作品、作者、机构、期刊和主题构成的庞大异质网络。传统科学计量学将其简化为孤立的表格指标(h指数、影响因子),这些指标忽略了拓扑上下文,且无法捕捉协同的不端行为。基于我们提出出版完整性图分析的配套综述,本文实现了该方法。我们在OpenAlex开放数据上定义了一个异质多变量图模型(7种节点类型、7种边类型),以及基于投影(引用和合著网络)、可解释结构指标、社区检测和三种异常出版模式筛选检测器的方法。我们刻意避免二分类:检测器返回带明确结构证据的候选排序结果,供人工评估。针对俄斯特拉发技术大学VSB的机构语料库(2020-2025年)及其单跳引用邻域,社区检测可恢复真实研究群体,中心性识别跨学科桥梁,筛选标记出密集合著派系、局部闭合引用环和主题孤立期刊。针对第二个以期刊为中心的语料库(含Scopus和DOAJ剔除的期刊作为外部真值,及规模匹配的对照组),朴素的病例对照设计产生看似强但虚假的检测器(存在突出性混淆),而匹配后唯一稳健信号是学科范围的广度(AUC为0.70);一种基于图的开放声望指标(期刊引用网络上的PageRank)可追踪JIF代理,同时比基于计数的指标抗引用操纵能力高一个数量级。我们将该方法发布为开源库apnet,具备可复现的CLI工作流和Web界面,分析可在普通硬件上数分钟内完成。
英文摘要
The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Building on our companion review, which proposed graph analysis of publishing integrity, this paper implements that approach. We define a heterogeneous multivariate graph model over OpenAlex open data (seven node types, seven edge types) and a methodology based on projections (citation and co-authorship networks), interpretable structural metrics, community detection, and three screening detectors of anomalous publishing patterns. We deliberately avoid binary classification: detectors return ranked candidates with explicit structural evidence for human assessment. On the institutional corpus of VSB - Technical University of Ostrava (2020-2025) with its one-hop citation neighbourhood, community detection recovers real research groups, centralities identify cross-disciplinary bridges, and the screenings flag dense co-authorship cliques, locally closed citation loops, and thematically isolated venues. On a second, venue-centric corpus with external ground truth (journals delisted by Scopus and DOAJ) and size-matched controls, a naive case-control design yields seemingly strong but spurious detectors (a prominence confound), whereas after matching the only robust signal is the breadth of disciplinary scope (AUC 0.70); an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy while being an order of magnitude more resistant to citation gaming than count-based indicators. We release the method as the open-source library apnet with a reproducible CLI workflow and a web interface; the analysis runs on commodity hardware in minutes.