域自适应预训练增强水处理语义表示,用于大规模结构化文献挖掘
Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining
浏览论文内容
中文总结 AI 辅助
本研究提出域自适应预训练模型WaterBERT,通过大规模语料持续预训练,在分类、命名实体识别和关系抽取任务上表现最佳,并构建知识图谱和检索系统,实现高效大规模水处理文献挖掘。
中文摘要 AI 辅助
水处理研究正在迅速扩展,但该研究获得的许多知识仍分散在非结构化文献中。该领域仍缺乏一个专用的语言模型,能够高效捕获水处理特定领域语义,以支持大规模文献挖掘。在此,我们通过开发WaterBERT来解决这一问题,这是一种域自适应编码器模型,专为水处理文本的语义表示和结构化信息提取而设计。WaterBERT通过在包含约29.7亿个token的大规模水处理语料库上进行持续预训练而开发。基于WaterBERT的三个微调模型在下游任务上进行了系统评估,在通用和领域特定BERT模型中取得了最佳整体性能,多类处理过程分类的F1分数为90.12%,命名实体识别为79.50%,关系抽取为74.04%。除了这些基准任务外,我们进一步展示了WaterBERT在大规模文献处理中的优势。应用于5,144篇《环境科学与技术》文章时,WaterBERT-BERTopic识别出连贯、多样且领域特定的研究主题,无需预定义类别。基于WaterBERT,我们以远低于商业LLM的成本处理了693,211篇摘要,同时保持具有竞争力的提取性能,构建了结构化水处理知识图谱。随后,该知识图谱与词汇和密集检索集成,开发了水知识增强检索系统(WaterKERS),其相关性得分为77.7,显著优于基于文本的检索基线(54.7-64.5)。通过WaterBERT,本研究为水处理研究中的大规模信息处理和证据映射提供了紧凑且可扩展的语义基础。
英文摘要
Water treatment research is expanding rapidly, but much of the knowledge acquired from this research remains scattered across unstructured literature. The field still lacks a dedicated language model that can efficiently capture water treatment-specific domain semantics for large-scale literature mining. Here, we address this by developing WaterBERT, a domain-adapted encoder model designed for semantic representation and structured information extraction from water treatment texts. WaterBERT was developed by continual pretraining on a large-scale water treatment corpus comprising about 2.97 billion tokens. Three fine-tuned models based on WaterBERT were systematically evaluated on downstream tasks, achieving the best overall performance among general-purpose and domain-specific BERT models, with F1 scores of 90.12% for multiclass treatment process classification, 79.50% for named entity recognition, and 74.04% for relation extraction. Beyond these benchmark tasks, we further demonstrated WaterBERT's advantages for large-scale literature processing. Applied to 5,144 Environmental Science & Technology articles, WaterBERT-BERTopic identified coherent, diverse, and domain-specific research topics without predefined categories. Building on WaterBERT, we processed 693,211 abstracts at substantially lower cost than commercial LLMs while retaining competitive extraction performance to construct a structured water treatment knowledge graph. The knowledge graph was then integrated with lexical and dense retrieval to develop a Water Knowledge-Enhanced Retrieval System (WaterKERS), which achieved a relevance score of 77.7, substantially outperforming text-based retrieval baselines (54.7-64.5). Through WaterBERT, this study provides a compact and scalable semantic foundation for large-scale information processing and evidence mapping in water treatment research.
发表机构
- UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales(新南威尔士大学)
- School of Electrical Engineering and Computer Science, The University of Queensland(昆士兰大学)
- Microsoft Copilot Studio AI(微软Copilot Studio AI)
- Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania(宾夕法尼亚大学)
- Department of Civil Engineering, The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。