arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更大的模型何时有用?针对本体学习的大语言模型规模的对照研究

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

Hamed Babaei Giglou, Sören Auer, Jennifer D'Souza

arXiv 2608.31118首次发表:更新:

发表机构

TIB Leibniz Information Centre for Science and Technology; L3S Research Center; Leibniz University of Hannover(莱布尼茨科学与技术信息中心(TIB); L3S研究中心; 汉诺威莱布尼茨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对照评估13个LLM后发现,模型规模并非本体学习的充分选择标准,架构和系列影响更大,为LLM辅助本体工程提供实证指导。

AI 中文摘要

大语言模型(LLM)规模对本体学习(OL)性能的影响仍未得到充分表征。我们使用OntoLearner检索增强生成流水线,对来自Qwen3.5和Qwen3.6系列的密集型与混合专家(Mixture-of-Experts)变体,以及专有GPT发布变体共13个模型进行了对照评估。所有模型在术语分类、分类体系发现和非分类关系提取任务中,均使用相同的嵌入模型、检索配置、提示模板、解码设置、数据集和指标,覆盖四个生物医学、材料科学与工程领域的本体。在密集型Qwen3.5系列中,增加参数规模主要提升精度而非召回率,最大提升出现在90亿至270亿参数之间。但规模的影响并非单调或在任务与领域间均匀:密集型270亿参数模型在术语分类上大幅优于更大的稀疏模型,而更大的混合专家模型在分类体系发现上取得开放权重模型中的最强结果。非分类关系提取在所有模型规模下都存在困难,尤其针对材料数据科学本体。匹配的Qwen变体与专有GPT发布变体间的性能差异进一步表明,架构和模型系列可能超过名义参数规模的影响。这些发现表明,仅模型规模不足以作为OL的选择标准,并为可复现的LLM辅助本体工程提供了实证指导。

英文摘要

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.

Comments14 pages, 1 figure, and 5 tables. WOP 2026 workshop at ISWC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑