发表机构
University of Padova(帕多瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对科学出版物增长带来的信息抽取需求,提出TWIX端到端IE流水线,采用两阶段框架解决GutBrainIE基准的所有子任务,在评测中大幅优于基线且排名第一,有效提升了精确率和召回率。
AI 中文摘要
科学出版物的指数级增长催生了对自动信息抽取(IE)系统的需求,以支持知识发现。在此背景下,GutBrainIE基准评测了肠道-大脑轴领域的命名实体识别(NER)、命名实体识别与消歧(NERD)以及关系抽取(RE)系统。我们提出了信息抽取两阶段工作流(TWIX),这是一种端到端的IE流水线,包含三个相互关联的模块,每个模块均采用两阶段框架来解决GutBrainIE的所有四个子任务。在开发集和测试集上的评估表明,我们的方法大幅优于基线方法,同时在所有子任务的所有参赛提交中排名第一。这些结果表明,所提出的两阶段流水线在实际场景中有效提升了精确率和召回率。
英文摘要
The exponential growth of scientific publications calls for automatic Information Extraction (IE) systems to support knowledge discovery. In this context, the GutBrainIE benchmark evaluates Named Entity Recognition (NER), Named Entity Recognition and Disambiguation (NERD), and Relation Extraction (RE) systems in the gut-brain axis domain. We propose Two-stage Workflow for Information eXtraction (TWIX), an end-to-end IE pipeline featuring three interconnected modules, each leveraging a two-stage framework to solve all four GutBrainIE subtasks. Evaluation on the development and test sets shows that our method substantially outperforms the baseline by a wide margin, while also ranking first among all participant submissions across all subtasks. These results indicate that the proposed two-stage pipeline effectively improves both precision and recall in practical settings.
CommentsAccepted at CLEF 2026: the 17th Conference and Labs of the Evaluation Forum