发表机构
Heidelberg University(海德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了德语和英语的多级别简化医学语料库PlainMedScale,开展试点研究发现现有可读性指标无法跨层级泛化、SOTA开放权重LLM的Plain Language提示仍保留部分输入难度,并公开了代码与数据。
AI 中文摘要
我们推出PlainMedScale,这是一个主题对齐的医学语料库,涵盖德语和英语的四个可理解性级别,数据来源于MSD(专业版和消费者版)、PubMed、《Apotheken Umschau Einfache Sprache》以及NHS。这四个层级对应不同的交际功能——参考、解释、决策支持和获取,突破了以往语料库中专家与普通大众的二元对比。借助该对齐开展的两项试点研究显示,在两种语域上建立的许多可读性指标无法在整个梯度范围内泛化,且针对Plain Language提示的SOTA开放权重LLM仍会部分保留输入的难度。代码(this https URL)和数据(this https URL)均已公开。
英文摘要
We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions --- reference, explanation, decision support, and access --- and move beyond the binary expert--lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open-weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS-Uni-Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.
Commentsaccepted at KONVENS 2026