面向科学出版物记录中资助者名称消歧的多功能嵌入模型
Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records
浏览论文内容
中文总结 AI 辅助
本文提出一个多语言多功能资助者名称消歧框架,利用多任务学习微调嵌入模型,在匹配准确率上显著超越通用大模型,并构建相似性网络处理未索引名称,为科学文献分析提供可复用方案。
中文摘要 AI 辅助
理解研究资助的历史分配和分布,有助于我们认识科学研究在不同领域、机构和地区的支持情况。然而,由于缺乏全面的资助者名称消歧解决方案,大规模分析受到阻碍,因为资助者名称常常存在拼写变体、翻译、缩写以及粒度不一致等问题。在本文中,我们提出了一个用于开发多语言、多功能资助者名称消歧模型的框架,并将其应用于生物多样性保护领域的研究出版物。为了构建训练数据集,我们将研究组织注册中心(ROR)(该中心为研究组织提供唯一标识符)与两个出版物数据集(Web of Science(WoS)和 Crossref 开放资助者注册中心(OFR))进行了整合。我们采用多任务学习,结合对比损失(Contrastive Loss)和多负样本排序损失(Multiple Negatives Ranking Loss),对来自 Sentence Transformer、Gemma 和 Qwen3 系列的三个开放权重嵌入模型进行了微调。性能最佳的模型在将 WoS 资助者名称与 ROR 标识符匹配时,准确率超过 0.90,比通用大语言模型(包括 GPT-5.2、Claude-Sonnet-4.6 和 Gemini-2.5-Flash)高出 0.1 以上。对于未在 ROR 中索引的资助者名称,我们构建了资助者名称之间的相似性网络,并识别了其中的聚类。最后,我们分析了消歧结果,并强调了因对较小资助者及非英语国家资助者了解有限而产生的挑战。这项工作为资助者名称消歧提供了一个可复用的框架,具有跨不同模型架构和数据集应用的潜力,其特点包括经济高效的训练数据创建以及多任务学习和消歧。
英文摘要
Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the lack of comprehensive funder name disambiguation solutions, as funder names often exhibit spelling variations, translations, abbreviations, and inconsistent levels of granularity. In this paper, we present a framework for developing multilingual, multi-functional funder name disambiguation models and demonstrate its application to research publications in biodiversity conservation. To construct a training dataset, we integrated the Research Organization Registry (ROR), which provides unique identifiers for research organizations, with two publication datasets: the Web of Science (WoS) and the Crossref Open Funder Registry (OFR). We used multi-task learning with Contrastive Loss and Multiple Negatives Ranking Loss to fine-tune three open-weight embedding models from the Sentence Transformer, Gemma, and Qwen3 families. The best-performing models achieved accuracy above 0.90 when matching WoS funder names to ROR identifiers, outperforming general-purpose LLMs, including GPT-5.2, Claude-Sonnet-4.6, and Gemini-2.5-Flash, by more than 0.1. For funder names not indexed in ROR, we constructed a similarity network among funder names and identified clusters within it. Finally, we analyzed the disambiguation results and highlighted challenges arising from limited knowledge of smaller funders and funders from non-English-speaking countries. This work provides a reusable framework for funder name disambiguation with potential applicability across different model architectures and datasets, featuring cost-effective training data creation and multi-task learning and disambiguation.
发表机构
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of Michigan(密歇根大学)
- Technical University of Munich(慕尼黑工业大学)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- Munich Data Science Institute(慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。