谁在使用开放权重模型?中国与科学领域AI的地理格局变迁
Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science
浏览论文内容
中文总结 AI 辅助
该研究分析2100万篇论文发现,GPT系列主导AI研究,开放权重模型使用量升至2026年的44.0%,中国机构研究人员采用率是其他机构的2.23倍,推动了开放权重模型的采用,反映AI模型生态的地缘文化重组。
中文摘要 AI 辅助
随着大型语言模型(LLM)成为科学研究的焦点,计算机科学家与科学技术研究(STS)学者已倡导使用开放权重模型。鉴于LLM研究已趋于成熟且有更多高质量模型系列可供选择,研究人员是否采用了开放权重模型?我们开展了首个针对科学研究中模型选择的系统性研究,分析了来自Semantic Scholar开放研究语料库(S2ORC)截至2026年6月的2100万篇全文文章。我们采用混合自然语言处理(NLP)流程提取文章全文中模型的出现情况,并确定研究人员是使用了这些模型还是仅提及它们。我们将语料库划分为单模型系列研究与多模型系列研究,将其分别作为应用AI研究与基础AI研究的替代指标。我们发现GPT系列模型在单模型系列与多模型系列研究中均占据主导地位,但这两个领域随时间推移正变得更加多样化。在单模型系列论文中,开放权重模型的使用稳步上升,到2026年达到44.0%。不过,我们发现近期的增长是由高质量开放权重中文模型的可用性推动的。进一步的逻辑回归模型显示,开放权重模型的采用呈异质性分布,经估算,中国机构的研究人员使用开放权重模型的几率是其他机构的2.23倍,这解释了2023年以来开放权重模型采用量增长的44.0%。补充的多项逻辑模型显示,这种关联集中在中文开放权重模型上:2026年,在带有中国机构隶属关系的论文中,其调整后使用率为37.1%,较无明显中国关联的论文高出27.9个百分点。这些发现表明,科学领域中开放权重模型的采用并非是向开放科学的普遍转变,而是模型生态系统更广泛重组的一部分,在此过程中,平台与市场以及塑造它们的社会文化和地缘政治背景,决定了哪些AI系统会成为科学工具。
英文摘要
As LLMs have become a flashpoint for scientific research, computer scientists and STS scholars have advocated the use of open-weight models. Since LLM research has matured and more high-quality model families are available, have researchers adopted open-weight models? We present the first systematic study of model selection in scientific research, analyzing 21 million full-text articles through June 2026 from the Semantic Scholar Open Research Corpus (S2ORC). We employ a mixed NLP pipeline to extract model occurrences in article full text and determine whether they are used or merely mentioned by researchers. We divide our corpus into single- and multi-model family studies, which we take as a proxy for applied and foundational AI research. We find GPT-family models dominate both single- and multi-family research, but that both areas are becoming more diverse over time. In single-family papers, open-weight model use rises steadily, reaching 44.0% in 2026. However, we find that recent growth is driven by the availability of high-quality open-weight Chinese models. Further, a logistic regression model finds that open-weight adoption is heterogeneously distributed, estimating that researchers at Chinese institutions have 2.23 times the odds of using an open-weight model, accounting for 44.0% of the increase in open-weight adoption since 2023. A complementary multinomial model shows this association is concentrated in Chinese open-weight models: in 2026, their adjusted use is 37.1% among papers with Chinese affiliations, a 27.9 percentage-point over papers with no observed China link. These findings suggest that open-weight adoption in science is not a general turn toward open science, but part of a broader realignment of model ecosystems in which platforms and markets, and the sociocultural and geopolitical contexts which shape them, determine which AI systems become scientific instruments.