arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型文学语料库的自动主题索引:一种适用于伏尔泰全集的机器学习方法

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

Miguel Arana-Catania, Gillian Pink, Glenn Roe

arXiv 2607.09316首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文以伏尔泰全集子语料库为测试案例,将自动主题索引任务设为多标签分类问题,比较多种机器学习方法,最佳模型F1分数达0.67,还评估跨语料库泛化及分析模型行为,为大规模文学和历史语料库的结构化主题访问提供了启示。

AI 中文摘要

主题索引,即将结构化概念标签分配给文本各部分的做法,对大规模文学和历史版本的学术访问至关重要,但仍主要是人工且劳动密集型过程。本文以伏尔泰《全集》的两个重要子语料库《论各民族的风俗与精神》和《百科全书问题》为测试案例,探讨机器学习在自动主题索引中的应用。该任务被构建为多标签分类问题,模型要为给定文本页面分配专业索引员会使用的索引条目集。我们比较了一系列方法,从带分类头的基于编码器的模型到通过低秩适应(LoRA)微调的生成式大语言模型(LLMs),模型规模从约30亿到1200亿参数不等。我们表现最佳的模型是4位量化配置的米斯特拉尔家族模型,F1分数高达0.67。鉴于专业索引的固有主观性以及模型预测虽与印刷索引不同但语义有效的频率,我们认为这些数字代表下限。我们还评估了跨语料库泛化,并对源文本中特别难以自动处理的文学和修辞特征的模型行为进行了详细的定性分析。我们的发现对为大规模文学和历史语料库提供结构化主题访问这一更广泛挑战具有启示意义。

英文摘要

Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: the Essai sur les mœurs et l'esprit des nations and the Questions sur l'Encyclopédie. The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given page of text. We compare a range of approaches -- from encoder-based models with classification heads to generative large language models (LLMs) fine-tuned via Low-Rank Adaptation (LoRA) -- spanning model sizes from approximately 3 to 120 billion parameters. Our best-performing model, from the Mistral family in a 4-bit quantised configuration, achieves F1 scores of up to 0.67; we argue that these figures represent lower bounds, given the inherent subjectivity of professional indexing and the frequency with which model predictions prove semantically valid despite diverging from the print index. We further evaluate cross-corpus generalisation and conduct a detailed qualitative analysis of model behaviour on literary and rhetorical features of the source texts that prove particularly resistant to automated treatment. Our findings have implications for the broader challenge of providing structured thematic access to large-scale literary and historical corpora.

Comments22 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑