arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COCI:会议组织者与内容标识符

COCI: Conference Organisers and Content Identifier

Angelo Salatino, Francesco Osborne, Alexis Vizcaino, Aliaksandr Birukou, Enrico Motta

arXiv 2608.24559首次发表:更新:

发表机构

Knowledge Media Institute, The Open University, UK

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对灰色文献中征稿通知与学术知识图谱孤立的问题,提出基于AI的COCI框架,结合LLMs与语义映射技术提取CfP结构化元数据,整合至知识库,为非出版商学术活动分析奠定基础。

AI 中文摘要

尽管灰色文献在学术交流中发挥着关键作用,但征稿通知(CfPs)等文献仍在很大程度上与现代学术知识图谱相互孤立。这些文档的非结构化和高度异质性特征,长期以来阻碍了对其进行大规模处理。在这篇演示论文中,我们提出了会议组织者与内容标识符(COCI),这是一种基于AI的框架,旨在从原始CfP文本中提取细粒度的结构化元数据。COCI采用多阶段流程,将大型语言模型(LLMs)与语义映射技术相结合,将提取的实体与已有的知识库(包括OpenAlex、DBLP、TIB ConfIDent和AIDA Dashboard)进行整合。通过对作者进行消歧,并将主题与会议系列进行语义对齐,COCI弥合了非正式学术传播与结构化语义网资源之间的差距,为系统分析非出版商相关学术活动奠定了基础。

英文摘要

Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated from modern Scholarly Knowledge Graphs. The unstructured and highly heterogeneous nature of these documents has traditionally hindered their large-scale processing. In this demo paper, we present the Conference Organisers and Content Identifier (COCI), an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts. COCI employs a multi-stage pipeline that combines Large Language Models (LLMs) with semantic mapping techniques to integrate extracted entities with established knowledge bases, including OpenAlex, DBLP, TIB ConfIDent, and the AIDA Dashboard. By disambiguating authors and semantically aligning topics and conference series, COCI bridges the gap between informal scholarly dissemination and structured Semantic Web resources, laying the foundation for systematic analysis of non-publisher-based academic events.

CommentsDemo paper accepted at ISWC 2026. To be presented in October 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑