arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OncoNoteBERT:面向真实世界门诊肿瘤学笔记自然语言处理的基础表示模型

OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes

Wuraola Oyewusi, Eliana Vasquez Osorio, Goran Nenadic, Gareth Price

arXiv 2610.03829首次发表:更新:

发表机构

The University of Manchester; The Christie NHS Foundation Trust(曼彻斯特大学; 克里斯蒂NHS基金会信托)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用临床模型难以处理肿瘤学笔记的问题,提出肿瘤学专用BERT模型OncoNoteBERT,结合持续预训练与定制分词,在困惑度和掩码预测上表现更优,强调表示层设计的重要性。

AI 中文摘要

真实世界的门诊肿瘤学笔记包含专业术语、肿瘤分期表达、治疗名称、毒性描述以及机构特定的去标识化标记,这些内容可能无法被通用生物医学或相邻领域的临床语言模型高效表示。我们利用一个受治理的英国门诊肿瘤学语料库(包含来自21,564名接受肺癌和头颈癌治疗患者的290,026份笔记)开发并评估了肿瘤学专用的BERT风格编码器。我们将RadBERT和PathologyBERT与两种本地策略进行了比较:OncoNote-RadBERT(通过持续掩码语言模型预训练生成)和OncoNoteBERT(使用肿瘤学WordPiece分词器从头训练)。模型通过验证集上的掩码语言建模损失和困惑度、分词器碎片化指标、临床术语分词、掩码标记探针以及探索性表示分析进行评估。两个外部编码器在零样本评估中对肿瘤学语料库拟合不佳(RadBERT困惑度为113.04;PathologyBERT为2035.03),而持续预训练产生了最强的拟合(OncoNote-RadBERT为2.10)。OncoNoteBERT实现了2.83的困惑度,但产生了最高效的分词,具有更低的子词繁殖率和更短的归一化序列长度。它还在13个掩码标记探针中的12个返回了临床可接受的预测,而OncoNote-RadBERT为13个中的7个。这种语料库级拟合与掩码标记性能之间的差异部分归因于分词器碎片化,而非仅学习到的语义。两个本地开发的模型都将机构占位符表示为单个可学习标记。这些发现表明,持续适应和定制分词提供了互补的益处,并且在将相邻领域编码器应用于肿瘤学NLP之前,表示层的设计至关重要。

英文摘要

Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.

CommentsAccepted at the AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models Workshop at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑