arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06167cs.AIcs.CL

基于生成式AI的模式引导分层信息提取与语义评估

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出模式引导的生成式AI框架,可零样本提取分层结构化信息并自动语义评估,在NICE文档上F1超90%,速度为人类专家30倍,具通用性与可迁移性。

中文摘要 AI 辅助

我们提出了一种基于模式的框架,用于使用生成式AI从非结构化文本文档中提取复杂的结构化信息,随后针对提取的信息与黄金标准进行自动语义评估。该模式作为编码领域知识的信息模型,为提取具有可变基数属性的分层嵌套信息及后续结果评估提供了统一、系统且一致的框架。从文档中提取信息以零样本模式单次调用模型完成。在评估步骤中,我们引入了一种基于路径的语义匹配算法,用于对齐提取结果与黄金标准中嵌套的可变基数属性;同时使用生成式AI对属性的提取值与黄金标准值进行语义比较,并引入 rubric(评分标准),根据领域特定考量将比较结果分类为完全匹配、语义匹配、有用匹配或不匹配。我们使用生成式AI模型Claude Opus 3,从英国国家卫生与临床优化研究所(NICE)发布的文档中成功提取了14个属性中的12个,F1分数超过90%;从单篇文档中提取属性所需时间约为人类领域专家的1/30。我们进一步验证了该框架在不同生成式AI模型间的通用性,以及在不同卫生技术评估(HTA)机构和语言间的可迁移性。

英文摘要

We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.

补充信息

↑