面向异构文档知识图谱构建的本体引导、去重感知抽取层
An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
浏览论文内容
中文总结 AI 辅助
本文提出一套本体引导的去重感知抽取层,用Qwen3.5-9B模型从异构文档抽取实体与关系,经五阶段优化后,在情报语料上使搜索召回率从70%升至95%且修正多类质量缺陷。
中文摘要 AI 辅助
大语言模型可流畅从非结构化文档中抽取实体与关系,但存在不一致问题:不同文档间类型词汇碎片化,同一人物有多个名称变体,关系重复,同名不同个体存在隐式合并风险。本文提出并设计、实现、实证优化了一套生产级抽取层,可将实时文档流转换为与形式本体对齐的经验证知识图谱。该系统从Kafka接收文档元数据,通过为各格式构建的处理器路由PDF、电子表格、Office及图像内容,使用经本体微调的本地部署Qwen3.5-9B模型分两轮抽取实体与关系。其核心特色是本体引导抽取:通过嵌入相似度从图数据库实时检索精选本体的相关片段并注入抽取提示,相比静态领域片段,可降低约94%的目录开销。抽取结果随后经过五阶段优化流程:确定性清洗、跨块合并、关系二次抽取、六种无需模型推理的去重算法,以及嵌入解析子系统(其冲突防护模块不受任何相似度分数影响)。在情报语料上的评估显示,搜索召回率从约70%提升至95%且无错误合并,还修正了七类隐式质量缺陷,涵盖从截断源文本单个字符的错误,到带头衔前缀实体的系统性重复等问题。
英文摘要
Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation. This paper presents the design, implementation, and empirical refinement of a production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology. The system consumes document metadata from Kafka, routes PDF, spreadsheet, Office, and image content through handlers built for each format, and extracts entities and relationships in two passes using a locally hosted Qwen3.5-9B model tuned on the ontology. Its distinguishing component is ontology-guided extraction: the relevant slice of a curated ontology is retrieved live from a graph database by embedding similarity and injected into the extraction prompt, reducing catalog overhead by about 94 percent relative to static domain slices. Extracted results then pass through a refinement pipeline of five stages: deterministic cleaning, merging across chunks, a second pass for relationships, six deduplication algorithms that require no model inference, and an embedding resolution subsystem whose conflict guard no similarity score can override. Evaluation on intelligence corpora improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect, ranging from a bug that truncated source text by a single character to the systematic duplication of entities that carried title prefixes.
发表机构
- Konectu(科内克图)
机构由 AI 辅助整理,请以论文原文为准。