arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于统一生成式训练框架的可跨域事件抽取系统

A Scalable Cross-Domain Event Extraction System via a Unified Generative Training Framework

Siting Liang, Omar Adjali, Omair Shahzad Bhatti, Daniel Sonntag

arXiv 2608.23261首次发表:更新:

发表机构

German Research Center for Artificial Intelligence; Carl von Ossietzky University of Oldenburg(德国人工智能研究中心; 奥尔登堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有事件抽取方法可扩展性与跨域泛化不足的问题,提出统一生成式序列到序列框架,微调预训练语言模型,构建支持多功能的网络应用平台,实现跨域事件抽取。

AI 中文摘要

事件抽取是信息抽取的基础。现有方法常将事件检测与论元抽取分离,或依赖特定数据集设计,限制了可扩展性与跨域泛化能力。我们提出一种统一生成式序列到序列框架,可联合执行事件抽取子任务,支持流水线与端到端两种配置。我们在多个不同领域的事件数据集上微调预训练语言模型,使单个模型既能保留领域特定语义,又能在庞大且不断演变的标签空间上实现泛化。我们通过为研究者与从业者定制的基于网络的应用平台验证了这些能力,该平台支持文档上传、感知事件抽取的模式、触发词与论元的可视化,以及不同领域间不同抽取配置的对比。

英文摘要

Event extraction is fundamental to information extraction. Prior approaches often separate event detection and argument extraction or depend on dataset-specific designs, limiting scalability and cross-domain generalization. We propose a unified generative sequence-to-sequence framework that performs event extraction subtasks jointly and supports both pipeline and end-to-end configurations. We fine-tune pretrained language models on multiple event datasets across diverse domains, enabling a single model to retain domain-specific semantics while generalizing over large and evolving label spaces. We demonstrate these capabilities through a web-based application tailored for researchers and practitioners. The platform supports document upload, schema-aware event extraction, visualization of triggers and arguments, and comparison of different extraction configurations across domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑