SCOPE和SCION:用于文本模式归纳与融合的基准和可审计参考管道
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text
浏览论文内容
中文总结 AI 辅助
研究针对模式图这一信息提取瓶颈,提出SCOPE基准和SCION参考管道。通过从训练文本构建候选空间,经严格JSON合同进行相关操作。在SCOPE核心套件上,SCION-lite表现最佳,SCION-RL减少对专有工程师依赖,还给出标准化结果及相关发布内容。
中文摘要 AI 辅助
模式图是基于模式的信息提取和知识图谱构建的上游瓶颈,但大多数提取系统都假定模式已可用。我们引入了SCOPE,这是一个仅基于训练文本的基准,用于从原始文本进行语料库到模式的归纳以及可选的模式融合。还介绍了SCION,它是一个可审计的参考管道。在SCOPE核心套件上,SCION-lite在多种指标下F1最高,SCION-RL变体减少了对专有LLM模式工程师的依赖。结果是针对标准化类型边缘目标报告的,发布内容包括证据关联输出等。
英文摘要
Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, built from 24 public information extraction sources (15 RE and 9 EE) normalized into evaluation-only gold schema graphs; its core event-extraction target covers event types and within-event argument roles, with inter-event links reported separately. We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it constructs candidate spaces from train text and restricts naming, merging, filtering, validation, and conservative fusion to candidate-linked evidence under strict JSON contracts. On the SCOPE core suite, SCION-lite attains the highest F1 among released source-schema references, Text2Onto-style, LLM-only, and matched extract-then-aggregate baselines under Literal, Fuzzy, Continuous, and Graph schema-graph metrics, while the compact open-model SCION-RL variant reduces reliance on proprietary LLM schema engineers. These results are reported against normalized typed-edge targets rather than as claims that induced schemas surpass human ontology design; the release includes evidence-linked outputs, parse/fallback logs, candidate retention/merging logs, run manifests, code, and benchmark packages at https://github.com/wandugu/paper_scion.
发表机构
- Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(信息工程研究所,中国科学院,北京,中国)
- School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(网络安全学院,中国科学院大学,北京,中国)
- School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国)
机构由 AI 辅助整理,请以论文原文为准。