MedDeID:基于真实或合成训练数据实现本地化管理的临床文本去标识化
MedDeID enables locally governed clinical-text de-identification from real or synthetic training data
浏览论文内容
中文总结 AI 辅助
MedDeID是一个本地化临床文本去标识化框架,利用内部标注与合成数据训练模型,在荷兰医院和初级保健基准上实现高召回率,并成功迁移至英语合成基准。
中文摘要 AI 辅助
临床笔记包含个人身份信息(PII),限制了其在研究和医疗AI中的重用,尤其是当数据不能离开机构时。我们开发了MedDeID,一个本地化框架,结合了内部标注和合成笔记生成,以及模型训练、推理、假名化和评估。在一个独立标注、经裁决的300份荷兰医院笔记基准上,医院训练的紧凑型变压器检测到98.9%的标识文本,同时仅删除了标注标识符之外的0.24%文本;仅用合成数据的对应模型检测到96.1%。在100份初级保健笔记上,合成训练的模型实现了比医院训练模型更高的召回率(90.3%对比87.0%),并对标识符格式扰动具有更强的鲁棒性。一个不使用真实文本训练的英语实例在两个外部合成基准上检测到99.7%和98.9%的标注标识符字符。这些结果证明了工作流到另一种语言的迁移,但不代表临床英语性能。MedDeID提供了一条使用真实或合成训练数据进行本地化管理的去标识化途径。
英文摘要
Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation and evaluation. On an independently annotated, adjudicated 300-note Dutch hospital benchmark, a hospital-trained compact transformer detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers; a synthetic-only counterpart detected 96.1%. On 100 primary-care notes, the synthetic-trained model achieved higher recall than the hospital-trained model (90.3% versus 87.0%) and greater robustness to identifier-format perturbations. An English instantiation trained without real text detected 99.7% and 98.9% of annotated identifier characters on two external synthetic benchmarks. These results demonstrate transfer of the workflow to another language, but not clinical English performance. MedDeID provides a route to locally governed de-identification using real or synthetic training data.
发表机构
- University of Antwerp(安特卫普大学)
- Antwerp University Hospital (UZA)(安特卫普大学医院)
机构由 AI 辅助整理,请以论文原文为准。