arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18927cs.AI

OntoBook:用于医学编码器预训练的基于本体的合成教科书

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

Rian Touchent, Éric de la Clergerie

中文总结 AI 辅助

研究提出OntoBook方法,将医学本体结构转为预训练信号,经随机游走、大语言模型重述等三阶段,用生成文本训练法语编码器,在多基准测试中有显著改进,强调目标对齐必要,还发布了相关教科书与模型检查点。

中文摘要 AI 辅助

我们提出了OntoBook,一种将医学本体结构转换为编码器语言模型预训练信号的方法。我们的方法有三个阶段:通过本体图的随机游走捕获医学代码之间的层次和因果关系,大语言模型将这些游走重新表述为流畅的教科书式散文,然后使用生成的文本在相同数据上以两个目标训练1.49亿参数的法语编码器ModernCamemBERT:掩码语言建模和代码对之间的关系预测。在三个法语医学编码基准上,OntoBook比仅进行掩码语言建模的预训练有显著改进。我们发现目标之间的对齐是必要的。我们发布了跨越三种法语本体的130万本大语言模型重新表述的医学教科书和预训练模型检查点。

英文摘要

We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through ontology graphs capture hierarchical and causal relations between medical codes, a large language model reformulates these walks into fluent textbook-style prose, and the resulting text is used to train ModernCamemBERT, a 149M-parameter French encoder, with two objectives on the same data: masked language modeling and relation prediction between code pairs. On three French medical coding benchmarks (FRACCO, Cantemist-FR, Distemist-FR), OntoBook achieves significant improvements over MLM-only pretraining, with +2.5 micro-F1 on FRACCO and +8.0 micro-F1 on Distemist. We find that alignment between objectives is necessary: misaligned training, where each task uses different data, causes a 30-point degradation. We release 1.3 million LLM-reformulated medical textbooks across three French ontologies (CIM-10, CCAM, ATC) and pretrained model checkpoints.

补充信息

↑