ROBE:用于从历史文本中提取极长尾事件的逆序偏倚专家模型
ROBE: Reversed-Order-Biased-Experts for Extracting Extreme Long-tail Events from Historical Texts
- Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
- Huygens Instituut(惠更斯研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对未被大模型预训练覆盖的19世纪前荷兰历史语料库,提出ROBE逆序偏倚专家模型,结合领域合成数据提升,使长尾事件提取的召回、精确率及F1值均优于简单微调编码器模型。
中文摘要 AI 辅助
本文提出了从涵盖17至18世纪的荷兰历史语料库中提取50余种事件的方法,旨在攻克“从长尾的长尾中提取信息”这一难题。19世纪前的历史数据本身属于小众领域,未被大型语言模型的预训练覆盖,而我们的目标是提取该领域可用训练数据中仅极少量标注的事件。我们针对训练数据中存在的事件子组创建专家分类器,分组依据为训练数据中的相似频率或语义相关性;对针对代表性不足事件训练的专家,赋予其更高的预测优先级,以避免被频率偏倚主导。这种专门用于保护长尾的分类器组合新方法被命名为ROBE:Reversed-Order-Biased-Experts(逆序偏倚专家模型)。我们还提出了一种创建领域特定合成数据的可控方法。我们的两种ROBE实现,相比简单微调的编码器模型,召回率分别提升0.10、精确率分别提升0.16;在我们的小众数据集的一组长尾类别中,最优模型的F1值提升0.10。
英文摘要
This paper proposes methods to extract over 50 types of events from a Dutch historical corpus spanning the 17th and 18th centuries. The methods we propose aim to tackle a very challenging scenario in Machine Learning: extracting the long-tail of the long-tail. Historic data from before the 19th century is in itself a niche domain not covered in the pre-training of Large Language Models, and we aim to extract events only scarcely annotated in the training data available for this domain. We propose creating expert classifiers for subgroups of the events present in the training data. We make these groupings based on similar frequency in the training data or on semantic relatedness. Experts trained on underrepresented events are assigned higher priority when predicting to avoid being dominated by frequency biases. We refer to this new way of combining classifiers, specifically tailored to protect the long-tail, as ROBE: Reversed-Order-Biased-Experts. We also propose a controlled method to create domain-specific synthetic data. Our two implementations of ROBE outperform a simple fine-tuned encoder model with a .16 increase in precision and a .05 increase in recall respectively. The best model achieves a .11 increase in f1 for a group of long-tail classes in our niche data set.