Meddies-PII:临床去标识化中个人身份信息提取的多语言框架
Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification
浏览论文内容
中文总结 AI 辅助
本文提出Meddies-PII框架,构建百万级多语言合成临床文档数据集,训练BIOES分类器,在十五个外部基准上平均F1达0.827,显著优于现有系统,并开源全部资源以推动临床去标识化研究。
中文摘要 AI 辅助
临床去标识化依赖于准确识别个人身份信息(PII)。然而,手动标注的数据集构建成本高昂,而现有的合成替代方案往往对其生成过程提供的细节有限,或依赖于相对简单的合成策略。我们引入了Meddies-PII-Dataset,一个包含一百万份合成临床文档的语料库,涵盖十七种语言和九种PII标签。这些文档使用属性条件提示生成,并通过十三个确定性门控进行验证,以确保结构和标注的一致性。为了评估该数据集的实用性,我们训练了Meddies-PII-Model,一个BIOES令牌分类器,并将其与现有的PII提取系统使用精确匹配的实体级F1分数进行比较。Meddies-PII-Model在所有报告的基准测试中取得了评估系统中最高的性能,在十五个外部基准测试中平均F1为0.827,而最强基线的平均F1为0.658。论文被接收后,我们将公开发布数据集、基准套件、模型、生成框架和评估代码,以支持多语言临床去标识化的研究。
英文摘要
Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited details about their generation process or rely on relatively simple synthesis strategies. We introduce Meddies-PII-Dataset, a corpus of one million synthetic clinical documents spanning seventeen languages and nine PII labels. The documents are generated using attribute-conditioned prompts and validated through thirteen deterministic gates that enforce structural and annotation consistency. To evaluate the dataset's utility, we train Meddies-PII-Model, a BIOES token classifier, and compare it with existing PII extraction systems using exact-match entity-level F1. Meddies-PII-Model achieves the highest performance among the evaluated systems on all reported benchmarks, with a mean F1 of 0.827 across fifteen external benchmarks, compared with 0.658 for the strongest baseline. Upon acceptance, we will publicly release the dataset, benchmark suite, model, generation framework, and evaluation code to support research on multilingual clinical de-identification.
发表机构
- Meddies AI
机构由 AI 辅助整理,请以论文原文为准。