arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.07224cs.CLcs.AI

TIDE 2.0:一个开放的、模型无关的临床笔记键控去标识化引擎

TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes

  • Stanford Medicine(斯坦福医学院)

机构由 AI 辅助整理,请以论文原文为准。

Jose D. Posada, Somalee Datta, Priya Desai

AI总结:

TIDE 2.0 是一个开源的键控去标识化引擎,通过可互换识别器和密码学生成替代值实现临床笔记的安全去标识化,并保持日期间隔和跨文档链接,在两个语料库上达到高召回率和精确率。

AI中文摘要:

临床笔记记录了患者护理过程中的大部分信息,但在移除受保护的健康信息(PHI)之前,这些笔记不能用于研究。去标识化通常被视为一个检测问题。仅检测是不够的:删除操作会连同标识符一起剥离临床内容,日期空白化会破坏纵向分析所需的时间间隔,而在每次出现时分配新的随机替代值会破坏患者笔记之间的联系。我们提出了TIDE 2.0,一个采用MIT许可证的引擎,包含两个可分离的阶段:一个可互换的识别器和一个键控匿名化器。两者都在机构拥有的硬件上运行。替代值通过密码学方式生成,无需存储链接表。日期按每位患者偏移,且保持间隔不变;在给定键下,每个值在所有出现中都获得相同的替代值;使用新键生成的发布版本无法与早期发布版本关联。我们还发布了TIDE2-Sentry,一个从大型语言模型中提炼出的识别器。在两个来自不同机构的黄金标注语料库上,默认配置在域内达到了0.88的片段级召回率,在第二个机构的语料库上达到了0.77,精确率分别为0.88和0.87。我们报告了每个类别的召回率和精确率以及这些汇总数据。该引擎是开源的,识别器在受限研究使用协议下可用,因此机构可以在自己的环境中运行、检查和扩展两者。

英文摘要:

Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PHI) is removed. De-identification is often treated as a detection problem. Detection alone is not sufficient: redaction strips clinical content along with identifiers, date blanking destroys the temporal intervals needed for longitudinal analysis, and assigning a fresh random surrogate at each occurrence breaks links between a patient's notes. We present TIDE 2.0, an MIT-licensed engine with two separable stages: an interchangeable recognizer and a keyed anonymizer. Both run on hardware the institution owns. Surrogates are generated cryptographically with no stored linkage table. Dates shift by a per-patient, interval-preserving offset; each value receives the same surrogate across all occurrences under a given key; and a release produced under a new key cannot be linked to earlier releases. We also release TIDE2-Sentry, a recognizer distilled from a large language model. On two gold-annotated corpora from two institutions, the default configuration reached span-level recall of 0.88 in-domain and 0.77 on the second institution's corpus, at precision 0.88 and 0.87. We report recall and precision per category alongside these aggregates. The engine is open source, and the recognizer is available under a gated research-use agreement, so institutions can run, inspect and extend both within their own environments.

↑