发表机构
Southern Methodist University; UT Southwestern Medical Center(南卫理公会大学; 德克萨斯大学西南医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发了VERGE智能体工作流,通过验证-优化循环提取临床笔记中的早发性结直肠癌预警症状及家族史风险,在4033条标注数据上较基线方法提升了性能,仅需1.5%人工审核。
AI 中文摘要
早发性结直肠癌在年轻成年人中发病率不断上升,但该人群的预警症状尚无循证指南用于后续检测,且结构化就诊数据无法捕捉支持早期检测及指导后续检测所需的细节,包括症状持续时间、背景以及结直肠癌的既定风险因素家族史。本研究旨在开发并评估一种自动化方法,用于从自由文本临床笔记中提取六种预警症状及家族史风险状态。我们开发了VERGE,这是一种智能体工作流,其中使用检索增强生成(RAG)提出初始标签和证据,随后经过有界验证-优化循环,该循环会检查文本接地性和临床有效性,修正并重新检查声明直至问题解决或达到上限,并将未解决的声明升级供人工审核。VERGE在4033条临床医生标注的笔记-发现对数据集上,与单智能体基线、基于规则的临床语言处理基线及替代底层语言模型进行了对比评估。与单智能体基线相比,VERGE减少了假阳性发现,将精确率从0.764提升至0.849,马修斯相关系数(MCC)从0.681提升至0.730,在精确率-召回率权衡中实现了平衡增益,且大部分标记错误可自主解决,仅1.5%的声明需要人工审核。这些结果表明,有界的、基于验证的工作流可减少不必要的阳性发现,同时不牺牲检测真实病例的能力。该方法为开发更可靠、可信赖的临床语言处理工具提供了路径,以支持年轻患者的结直肠癌风险评估。
英文摘要
Early-onset colorectal cancer is increasing among younger adults, yet red-flag symptoms in this age group have no evidence-based guidelines for follow-up testing, and structured encounter data do not capture the detail needed to support early detection and inform follow-up, including symptom duration, context, and fam- ily history, an established colorectal-cancer risk factor. This study aimed to develop and evaluate an automated method for extracting six red-flag symptoms and family-history risk status from free-text clinical notes. We developed VERGE, an agentic workflow in which an initial label and evidence are proposed using retrieval-augmented generation, then passed through a bounded verification- refinement cycle that checks textual grounding and clinical validity, corrects and rechecks a claim until resolved or a limit is reached, and escalates unresolved claims for human review. VERGE was evaluated on 4,033 clinician-labeled note-finding pairs against a single-agent baseline, a rule-based clinical language-processing baseline, and an alternative underlying language model. Compared with the single-agent baseline, VERGE reduced false positive find- ings, improving precision from 0.764 to 0.849 and MCC from 0.681 to 0.730, a balanced gain across the precision-recall trade-off, and resolved most flagged errors autonomously, with human review required for only 1.5 percent of claims. These results indicate that a bounded, verification-based workflow can reduce unnecessary positive findings without sacrificing the ability to detect true ones. This approach offers a path toward more reliable and trustworthy clinical language-processing tools to support colorectal cancer risk assessment in younger patients.