arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

阅读新闻:通过持续预训练使大型语言模型适配瑞典新闻业

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

Lukas Borggren, Jenny Kunz, Marco Kuhlmann

arXiv 2608.30609首次发表:更新:

发表机构

Linköping University; Bonnier News(林雪平大学; 邦尼尔新闻)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过整理新闻数据集对大型语言模型进行持续预训练,结合经验回放缓解遗忘,提升了模型在瑞典新闻领域的生成质量与事实知识,且需针对性评估适配效果。

AI 中文摘要

大型语言模型的通用能力日益增强,但在小众或研究不足的领域,其实用性仍可能有限。解决这一局限的一种方法是,通过在目标领域语料库上进行额外训练,使现有模型专门化。本研究中,我们探究此类持续预训练方法,以适配大型语言模型到瑞典新闻业,使用的是我们从数百万篇新闻文章中整理的高质量数据集。为评估适配效果,我们还构建了一个新的领域特定基准,涵盖六项编辑任务。在两种模型规模上进行全量微调与参数高效微调后,我们发现持续预训练在目标领域能带来益处,但仅在搭配经验回放以缓解遗忘时才成立。我们观察到模型的生成质量与事实知识有持续提升,但判别式任务的熟练度并未提升。探索一种促进指令遵循的无训练方法后,我们看到了进一步的改进,但这仅针对使用低秩适配训练的模型。至关重要的是,我们证明了适配过程中针对性评估的重要性,因为现有瑞典基准在很大程度上无法捕捉模型的领域内性能提升。

英文摘要

Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional training on target-domain corpora. In this work, we investigate such continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that we curate from millions of news articles. To evaluate the adaptation efficacy, we also construct a novel domain-specific benchmark that covers six editorial tasks. Through full and parameter-efficient fine-tuning across two model sizes, we find that continued pre-training yields benefits in the target domain, but only when paired with experience replay to mitigate forgetting. We observe consistent enhancements in the models' generation quality and factual knowledge, but not their proficiency in discriminative tasks. Exploring a training-free method to facilitate instruction following, we see further improvements, but exclusively for models trained with low-rank adaptation. Crucially, we demonstrate the importance of targeted evaluation in the adaptation process, as an existing Swedish benchmark largely fails to capture the models' in-domain performance gains.

CommentsAccepted at EMNLP 2026 Industry Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑