arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Climate-ModernBERT:重新审视语料库构成以实现领域自适应持续预训练

Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining

Yongan Yu, Shantam Raj, Jingwei Ni, Ario Saeid Vaghefi, Dominik Stammbach, Markus Leippold

arXiv 2609.07798首次发表:更新:

发表机构

McGill University; University of Zürich; Princeton University; ETH Zürich; Mila – Quebec Artificial Intelligence Institute(麦吉尔大学; 苏黎世大学; 普林斯顿大学; 苏黎世联邦理工学院; 米拉-魁北克人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Climate-ModernBERT,通过持续预训练和参数空间合并整合异构气候语料,在九个基准上以76.3平均F1超越基线2.8点,证实学术语料适应信号最强。

AI 中文摘要

气候领域的自然语言处理(NLP)要求模型处理异构文本来源,包括科学文献、政策披露和合成报告。然而,在持续预训练(CPT)期间如何有效组合多样化的领域语料库仍未得到充分探索。我们提出了Climate-ModernBERT,这是一个通过在现代BERT基础模型上对三个气候语料库进行持续预训练而获得的气候适应型编码器模型家族:学术气候文本、气候过滤的网页数据和合成气候文档。我们系统地比较了在语料库混合上进行联合持续预训练与对独立专门化检查点进行参数空间合并的方法。在九个气候NLP基准测试中,我们最好的模型达到了76.3的平均F1分数,比普通现代BERT基线显著提高了2.8个百分点。在气候NLP设置中,结果表明学术气候语料库在评估的来源中提供了最强的适应信号,而参数空间合并优于联合多源训练,并且更好地保留了来自异构气候语料库的互补信息。我们发布了所有Climate-ModernBERT变体和训练检查点,以支持气候NLP和领域自适应预训练的未来研究。

英文摘要

Natural Language Processing (NLP) in the climate domain requires models to process heterogeneous text sources, including scientific literature, policy disclosures, and synthetic reports. However, how to effectively combine diverse domain corpora during continued pretraining (CPT) remains underexplored. We introduce Climate-ModernBERT, a family of climate-adapted encoder models obtained through continued pretraining of ModernBERT-Base on three climate corpora: academic climate text, climate-filtered web data, and synthetic climate documents. We systematically compare joint continued pretraining on corpus mixtures with parameter-space merging of independently specialized checkpoints. Across nine climate NLP benchmarks, our best model achieves 76.3 average F_1, improving significantly over a vanilla ModernBERT baseline by 2.8 points. Within the climate NLP setting, the results show that academic climate corpora provide the strongest adaptation signal among the evaluated sources, while parameter-space merging improves over joint multi-source training and better preserves complementary information from heterogeneous climate corpora. We release all Climate-ModernBERT variants and training checkpoints to support future research in climate NLP and domain-adaptive pretraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑