arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17051cs.CLcs.AI

基于机构特定的大语言模型提示工程恢复去标识化系统及其黄金标准均遗漏的受保护健康信息

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出校准良好的上下文学习(ICL)的大语言模型(LLM),可通过机构特定提示恢复去标识化系统遗漏的受保护健康信息(PHI),平衡精确率与召回率,是专用去标识化系统的有效替代方案。

中文摘要 AI 辅助

电子健康记录的二次使用需要去标识化,但现有系统会遗漏与机构场景相关的受保护健康信息(PHI),例如医院缩写、建筑名称和内部代码,这些信息的状态由本地决定。本文探究具有上下文学习(ICL)能力的大语言模型(LLM)是否能缩小这一差距并控制精确率-召回率的权衡。在来自德克萨斯儿童医院的100份标注儿科肿瘤病历(共5322个PHI片段)上,我们将8种LLM与2种专用系统(斯坦福TiDE、OpenMed PII)以及2种基于模式的基线进行了基准测试。每种LLM在三种特异性递增的提示下运行:(1)符合HIPAA的基线;(2)基线加上其遗漏的机构PHI类别;(3)提示2加上防止过度编辑临床内容的指令。随后我们将14种多智能体和集成配置与最佳单提示进行了比较,其中召回率为主要安全指标。LLM的表现优于专用系统(最佳F1值为0.918±0.001,而TiDE为0.779),优势集中在上下文类别上。明确遗漏的类别可恢复79%(48/61)的相关信息,而抑制过度编辑则能恢复精确率。没有任何智能体架构优于校准后的单遍提示(F1值为0.906-0.907),但LLM输出揭示了414个候选标注缺口;重新标注确认了227个PHI片段,针对这些片段,最终提示达到了0.981的召回率(F1值为0.907±0.002)。校准良好的ICL可在每份病历调用一次LLM的情况下,同时解决机构PHI缺口和精确率-召回率的权衡问题。LLM的运行成本高于传统方法,但该成本换来了一种审计参考标准的方式。LLM是专用去标识化系统的合法、适应性替代方案;机构特定的提示开发应作为主要适应策略。

英文摘要

Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.

发表机构

  • Baylor College of Medicine(贝勒医学院)
  • Jan and Dan Duncan Neurological Research Institute(简与丹·邓肯神经学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑