arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证用于自然语言生成的DBpedia三元组集合

Validating DBpedia Triple Sets for Natural Language Generation

Mark Andrade, Simon Mille, Anya Belz, Brian Davis

arXiv 2609.07589首次发表:更新:

发表机构

Dublin City University; University of Bordeaux; CNRS(都柏林城市大学; 波尔多大学; 法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究从自然语言生成角度评估DBpedia三元组质量,提出过滤可疑三元组的方法,在保持高精确率的同时提升召回率。

AI 中文摘要

我们从自然语言生成的角度研究了单个DBpedia三元组的质量,并提出并评估了一种收集特定实体三元组集合的方法,该方法在过滤掉可疑三元组的同时,最大限度地减少正确三元组的损失。我们在与手动标注数据的评估中表明,通过验证规则,三元组选择的精确率可以达到98%,并且通过对少数属性定义的改进,可以在不损害精确率的情况下将召回率提高40%。

英文摘要

We present a study of the quality of individual DBpedia triples from the perspective of Natural Language Generation, and propose and evaluate an approach for collecting entity-specific triple sets that filters out questionable triples while minimizing the loss of correct ones. We show in an evaluation against manually annotated data that with validation rules, it is possible to reach 98% precision in triple selection, and with improvements to a few Property definitions, it is possible to improve recall by 40% without harming precision.

Comments20 pages, 1 figure 18 tables, https://github.com/andradeM17/DBpedia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑