发表机构
Massachusetts General Hospital; University of Washington(麻省总医院; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对临床症状检测中信息难结构化及现有方法不足的问题,提出多智能体系统Pythia,无需人工提示工程或微调,能自主优化提取提示。通过与词汇表比较,验证其在临床记录症状提取上的有效性及推广性,优于部分传统方法。
AI 中文摘要
临床记录包含许多使患者就医的体征和症状,但这些信息很少进入结构化字段。现有提取方法要么依赖产生误报的上下文无关规则,要么依赖需要大量微调的监督模型。我们提出了Pythia,一个多智能体系统,它能自主编写和优化临床概念的提取提示,无需人工提示工程或微调。Pythia在本地托管的开放权重模型上运行,将临床记录保存在本地基础设施上,并根据开发集的敏感性和特异性选择提示。我们将Pythia与一个精心策划的词汇表在400份代表387名患者的临床记录中的72种体征和症状上进行了比较。每个概念的开发集(n = 300)和验证集(n = 100)独立划分。Pythia的平均敏感性为0.76,特异性为0.95,而词汇表分别为0.82和0.76,在62个直接可比概念中的20个概念上,Pythia在这两个指标上匹配或超过了词汇表。对于词汇表将每份记录都标记为阳性的14个概念,Pythia通过要求是现在时态、患者归因的发现而不是对术语的任何文本提及,恢复了0.97的平均特异性。特异性从开发集转移到验证集时,在不同患病率下退化最小,而敏感性转移在患病率低于5%时减弱,在患病率低于2%时平均差距达到0.25。在相同开发集上按每个概念微调的BERT分类器平均敏感性为0.23,对于患病率低于约5%的概念,敏感性降至零。这些发现表明,自主、无需微调的提示优化可以产生症状提取提示,能从开发集有效推广到验证集,同时仍可在本地基础设施上部署。
英文摘要
Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning. We present Pythia, a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, Pythia keeps clinical notes on local infrastructure and selects prompts using development-set sensitivity and specificity. We compared Pythia with a curated lexicon across 72 signs and symptoms from 400 clinical notes representing 387 patients. Development (n=300) and validation (n=100) sets were partitioned independently for each concept. Pythia achieved mean sensitivity of 0.76 and specificity of 0.95, compared with 0.82 and 0.76 for the lexicon, and matched or exceeded the lexicon on both metrics for 20 of 62 directly comparable concepts. For 14 concepts where the lexicon labeled every note positive, Pythia recovered mean specificity of 0.97 by requiring a present-tense, patient-attributed finding rather than any textual mention of a term. Specificity transferred from development to validation with minimal degradation across prevalences, whereas sensitivity transfer weakened below 5% prevalence, reaching a mean gap of 0.25 below 2% prevalence. A BERT classifier fine-tuned per concept on the same development set achieved mean sensitivity of 0.23 and collapsed to zero sensitivity for concepts below roughly 5% prevalence. These findings suggest that autonomous, fine-tuning-free prompt optimization can produce symptom extraction prompts that generalize effectively from development to validation while remaining deployable on local infrastructure.