基于本体扩展和越南历史教科书检索的自动树状知识图谱构建
Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks
- The University of Danang(岘港大学)
- Vietnam-Korea University of Information and Communication Technology(越南-韩国信息与通信技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对低资源语言越南语,提出端到端树状知识图谱构建与检索评估流水线,构建了含750节点的知识图谱,其自上而下图遍历策略在NDCG@10上超基线4.7个百分点。
AI中文摘要:
基于层次化知识图谱(KG)的检索增强生成(RAG)已成为为大语言模型提供结构化知识支持的强大方法,但存在两大核心挑战:其一,针对越南等低资源语言,缺乏利用本体扩展实现自动知识图谱构建的方法;其二,缺乏对利用层次结构的知识检索策略的系统性评估。本文提出一种用于知识图谱构建和检索策略评估的端到端流水线。在知识图谱构建中,采用三阶段混合关系抽取流水线:通过Union-Find实现批内去重、近似批间搜索,以及结合质心过滤器(减少提示词数量)和五步双大语言模型验证器的大语言模型抽取,以防止本体膨胀。两层架构由不可合并的结构节点(保留文档结构)和可合并的内容节点组成。检索评估包含三种图遍历策略:自上而下、水平和自下而上,在从109个子图生成的1210个越南语查询的合成基准上进行评估,这些查询按五种查询方向分类。本文从近400页的越南高中历史教科书中构建树状知识图谱,生成750个节点和4341条语义边,本体类型从40种控制增长至41种。在实验的图遍历策略中,带结构的自上而下策略在NDCG@10指标上超过向量基线4.7个百分点。结果表明,树状结构信息提供了超出平面余弦相似度的有价值信息,但当查询不需要结构上下文时会降低性能。
英文摘要:
Hierarchical Knowledge graph (KG)-based retrieval augmented generation (RAG) has emerged as a powerful approach for supporting large language models with structured knowledge. However, there are primary challenges: (i) the lack of methods for automatic KG construction using ontology expansion for low-resource languages such as Vietnamese, (ii) the absence of systematic evaluation for knowledge retrieval strategies leveraging the hierarchical structures. In this paper, we propose an end-to-end pipeline for KG construction and retrieval strategies evaluation. In the KG construction, we employ a three-phase hybrid relation extraction pipeline: intra-batch deduplication via Union-Find, approximate cross-batch search, and LLM extraction with a centroid filter that reduces prompts combined with a five-step dual-LLM validator to prevent bloated ontology. A two-tier architecture consists of unmergeable structural nodes to preserve the document structure and mergeable content nodes. The retrieval evaluation consists of three graph traversal strategies: Top-Down, Horizontal, and Bottom-Up, which are evaluated on a synthetically generated benchmark of 1,210 Vietnamese queries from 109 subgraphs, categorized by five query directions. In this paper, we construct the tree knowledge graph from Vietnamese high school History textbooks (nearly 400 pages) to produce 750 nodes and 4,341 semantic edges with controlled ontology growth from 40 to 41 types. Among experimental graph traversal strategies, the Top-Down strategy with structure surpasses the vector baseline by 4.7 percentage points in NDCG@10. As a result, tree-structural information provides valuable information beyond flat cosine similarity but degrades performance when the query does not require structural context.