AI 中文总结
该研究提出多检测器等价性评估协议,发现格式良好的同类型PHI代理替换不会显著影响下游PHI检测器的性能,且公开了评估资源供审计使用。
AI 中文摘要
保留结构的去标识化方法会将受保护健康信息(PHI)替换为真实的同类型代理,例如将“Anna S.”替换为“Maria S.”,而非替换为[NAME],以此保证临床文本的流畅性并让下游工具正常运行。但这种方法仅在替换本身未破坏下游工具依赖的信号时才有效。我们提出一个可检验的窄问题:在去标识化工具实际掩码的文本片段上,下游PHI检测器仍能检测到这些代理吗?我们引入了一种配对的多检测器评估协议,该协议:(i)仅在掩码片段上对效用评分,将覆盖范围与效用解耦;(ii)采用等价性检验(TOST)而非零假设显著性检验,后者在我们的样本量(5.7万对片段)下无参考价值;(iii)构建代理失效类型学,将可修复的生成器缺陷与检测器的内在限制区分开。在11种检测器、7个基准和7种语言(共1750篇文档)的实验中,掩码片段的召回率从76.1%降至74.9%——我们的等价性检验显示,在±2个百分点的范围内,该变化与零具有统计等价性(p≈3e-9),且检测器的排名保持不变。剩余的损失并非反映检测器检测PHI的能力下降,而是集中在畸形和分布外的代理上(例如截断的“Chicago”变为“Illino”,显著性损失的“Cedars-Sinai”变为“Vidant”)。编辑底线和开源代理基线表明,该效应是格式良好替换的属性,而非某一工具的属性。我们在该httpsURL发布了评估子集、评分代码和交互式仪表板,以便该协议可审计任何保留结构的变换。
英文摘要
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this only helps if the substitution does not itself corrupt the signal those tools rely on. We ask a narrow, testable question: on the spans a de-identifier actually masks, can downstream PHI detectors still find the surrogate? We introduce a paired, multi-detector evaluation protocol that (i) scores utility only on masked spans, decoupling coverage from utility; (ii) uses equivalence testing (TOST) rather than null-hypothesis significance testing, which is uninformative at our sample size (57k paired spans); and (iii) builds a surrogate-failure typology separating fixable generator defects from intrinsic detector limits. Across 11 detectors, 7 benchmarks, and 7 languages (1,750 documents), recall on masked spans moves from 76.1% to 74.9% -- a change our equivalence test shows is statistically equivalent to zero within a +/-2-point margin (p ~ 3e-9), with detector ranking preserved. The residual loss does not reflect detectors getting worse at PHI: it concentrates in malformed and out-of-distribution surrogates (truncation Chicago -> Illino, salience loss Cedars-Sinai -> Vidant). A redaction floor and an open-source surrogate baseline indicate the effect is a property of well-formed substitution, not of one tool. We release the evaluation subsets, scoring code, and an interactive dashboard at https://custodianai.pages.dev so the protocol can audit any structure-preserving transform.
Comments12 pages, 3 figures, 10 tables. Code, data, and interactive dashboard: https://custodianai.pages.dev ; repository: https://github.com/Custodian-Labs/guardian-layer-phi-benchmark