arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLMs4OL 2026任务的团队:旗舰任务与复用:面向本体学习的检索增强生成及词汇约束过滤

pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning

Shivam Mishra, Dhannu Ram Meena, Muneendra Ojha, Krishna Pratap Singh, Kuldeep Singh

arXiv 2608.27101首次发表:更新:

发表机构

Indian Institute of Information Technology Allahabad(印度阿拉哈巴德信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该团队针对LLMs4OL 2026挑战赛的两个本体学习任务,采用检索增强生成与词汇约束过滤方法,取得了特定指标结果,但存在未提取非分类关系的局限。

AI 中文摘要

尽管大型语言模型(LLMs)已取得显著进展,但从文本中学习本体仍面临挑战,这类模型会产生领域术语幻觉、输出格式不一致,且更倾向于层级关系而非关联关系。在LLMs4OL 2026挑战赛中,我们采用离线检索增强的少样本提示流程,同时应对端到端旗舰任务(任务A)与本体扩展复用任务(任务B)。我们的系统使用Qwen2.5-14B-Instruct模型搭配all-MiniLM-L6-v2进行示例检索,为任务A选择前5个示例,为任务B选择前2个示例;采用左截断上下文窗口策略,在长提示中保留任务指令。对于任务B,生成的三元组会经过确定性词汇约束过滤:当三元组至少一个端点属于样本的封闭术语/类型词汇时保留,同时移除初始本体的重复项。该方法在任务B上的语义图相似度为0.8692、术语分类F1值为0.9200、分类体系发现F1值为0.8540,任务A的语义图相似度为0.7416;但未提取到非分类关系,凸显了封闭、面向分类体系的关系词汇表的局限性。

英文摘要

Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain terms, produce inconsistent formats, and favor hierarchical over associative relations. In the LLMs4OL 2026 Challenge, we address both the End-to-End Flagship Task (Task A) and Ontology Extension Reuse Task (Task B) using an offline retrieval-augmented few-shot prompting pipeline. Our system employs Qwen2.5-14B-Instruct with all-MiniLM-L6-v2 for demonstration retrieval, selecting the top-5 examples for Task A and top-2 for Task B. A left-truncated context-windowing strategy preserves task instructions within long prompts. For Task B, generated triples undergo deterministic vocabulary-constrained filtering, retaining triples when at least one endpoint belongs to the sample's closed term/type vocabulary and removing duplicates of the initial ontology. The approach achieves Semantic Graph Similarity of 0.8692, Term-Typing F1 of 0.9200, and Taxonomy Discovery F1 of 0.8540 on Task B, while Task A achieves 0.7416 Semantic Graph Similarity. However, no non-taxonomic relations are extracted, highlighting limitations of closed, taxonomy-oriented relation vocabularies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑