生成式人工智能是否取代了监督式XMLC?关于德国科学文献自动主题索引的基准研究
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
浏览论文内容
中文总结 AI 辅助
研究德国科学文献自动主题索引任务,对比监督式XMLC方法、经典词汇匹配基线及基于大语言模型的方法,发现基于Transformer密集特征的监督式XMLC算法在二元相关性指标最佳,基于大语言模型的生成式方法在分级相关性及长尾性能上更佳。
中文摘要 AI 辅助
以大量受控词汇作为标签集,图书馆中的自动主题索引任务可理解为多标签分类任务。若主题词集很大,该问题符合极端多标签分类(XMLC)目标。本研究将一系列专业监督式XMLC方法应用于德国国家图书馆收集的当代德国科学文献主题索引测试案例。通过纳入经典词汇匹配基线和我们最近开发的三种基于大语言模型的方法进行基准对比。算法在多个指标上进行评估和比较,包括与先前索引材料的二元相关性比较以及专业主题馆员的分级相关性评级。所有方法面临的挑战是从主题词汇的长尾中可靠地提出建议。我们发现,基于Transformer的密集特征的监督式XMLC算法在整体二元相关性指标方面给出最佳结果。然而,在分级相关性和主题词汇长尾中的性能方面,基于大语言模型的生成式方法给出更好结果,使其成为未来实际应用的有前途的替代方案。
英文摘要
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.
发表机构
- Deutsche Nationalbibliothek(德国国家图书馆)
机构由 AI 辅助整理,请以论文原文为准。