发表机构
Sant'Anna School of Advanced Studies(圣安娜高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出端到端对账流水线,从软件期刊和可复现性报告中采集4,397对DOI-仓库链接,设计两个Wikidata应用配置文件,将学术记录与归档源代码连接,并创建4,182个新软件条目。
AI 中文摘要
软件是一流的科学对象,然而源代码与学术记录之间经过验证的链接在关联开放数据(LOD)云中仍然基本缺失,使得归档工件与语义发现相隔离。本文提出了一种端到端的对账流水线,从论文与其源代码之间链接明确且经编辑验证的来源中采集、验证并建模出版物-仓库对:这些来源包括以软件为中心的期刊 JOSS、SoftwareX 和 IPOL,以及 SIGMOD 可用性与可复现性倡议(ARI)的可复现性报告。这产生了包含 4,397 个 ⟨DOI,仓库URL⟩ 对的精选语料库。我们设计了两个基于 Wikidata 类(一个用于学术文章,一个用于软件实例)的独特应用配置文件,并与 this http URL 和 CodeMeta 词汇表对齐。这种架构分离使得规则基础的对账能够在两个粒度上进行:轻量级的内联出版物引用,或配备 SWHID(Software Heritage 的内容寻址标识符)的独立一流 Wikidata 软件节点。对 Wikidata 的只读查询显示,采集的仓库中仅有 82 个已在其中建模;经人工审核的批次随后创建了 4,182 个与其文章交叉链接的新软件条目。我们进一步表明,新兴的 COAR Notify 协议(我们未参与开发的外部工作)的负载能够原生映射到我们的输入格式,因此同一后端未来可服务于实时增强流。我们的核心贡献是一对应用配置文件,将 Wikidata 转变为学术记录与归档源代码之间的连接器;我们公开发布所有代码、应用配置文件和采集的数据集。
英文摘要
Software is a first-class scientific object, yet validated links between source code and the scholarly record remain largely absent from the Linked Open Data (LOD) cloud, isolating archived artefacts from semantic discovery. This paper presents an end-to-end reconciliation pipeline that harvests, validates, and models publication-to-repository pairs from sources where the link between a paper and its source code is explicit and editorially verified: the software-centric journals JOSS, SoftwareX, and IPOL, together with the reproducibility reports of the SIGMOD Availability and Reproducibility Initiative (ARI). This yields a curated corpus of 4,397 $\langle$DOI, repository-URL$\rangle$ pairs. We design two distinct application profiles grounded in Wikidata classes (one for scholarly articles, one for software instances) aligned with the schema.org and CodeMeta vocabularies. This architectural separation enables rule-based reconciliation at two granularities: lightweight, inline publication references or standalone, first-class Wikidata software nodes equipped with SWHIDs, Software Heritage's content-addressed identifiers. A read-only lookup against Wikidata shows that only 82 of the harvested repositories were already modelled there; human-reviewed batches have since created 4{,}182 new software items cross-linked to their articles. We further show that payloads of the emerging COAR Notify protocol, an external effort we do not develop, map natively onto our input format, so the same backend could later serve a live enrichment stream. Our core contribution is a pair of application profiles that turn Wikidata into a connector between the scholarly record and archived source code; we openly release all code, application profiles, and harvested datasets.
Comments15 pages, 4 figures. Accepted at the 7th Wikidata Workshop (Wikidata 2026), co-located with ISWC 2026. Open-source pipeline and code available at https://github.com/ftosoni/swh-wd-reconciliation