大型语言模型用于结构化临床数据分析:双智能体接地与验证
Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
浏览论文内容
中文总结 AI 辅助
提出CLEAR-Med双智能体框架,分离SQL调用与独立验证,实现结构化临床数据的可追溯自然语言分析,在21站点新生儿数据上较基线提升54.4个百分点。
中文摘要 AI 辅助
目的:开发并描述CLEAR-Med,一个用于结构化临床数据自然语言分析的双智能体框架,该框架将基于SQL的调用与独立验证分离。方法:CLEAR-Med使用一个智能体将问题转换为可执行的结构化查询语言(SQL),保留已执行的查询和数据库结果,并生成草稿。确定性检查和单独调用的跨提供者验证智能体随后接受草稿、请求一次有界修复或弃权(不执行)。我们将该系统形式化为一个有界选择性流水线,并在一个包含532条去标识化婴儿记录和约1300个变量的协调的21站点新生儿缺氧缺血性脑病表上,评估了CLEAR-Med的配置和可扩展性,以及调用智能体在25个查询开发基准上的准确性和一致性。结果:CLEAR-Med完成了所有六种名义可扩展性配置,包括500x1300。在25个开发基准查询重复五次中,调用智能体正确回答了125个响应中的83个(66.4%;查询簇自助95%置信区间,48.0-83.2%),而未接地的ChatGPT基线为125个中的15个(12.0%;95%置信区间,3.2-22.4%),配对改进为54.4个百分点(95%置信区间,36.8-72.0%)。结论:CLEAR-Med为结构化临床数据的可追溯分析提供了一种通用架构:数值声明保持与已执行的SQL关联,未解决的案例可以失败关闭。报告的实验表征了CLEAR-Med的配置和可扩展性以及调用智能体的准确性,而形式化分析建立了完整控制流的编码属性保证;对验证和弃权阶段的前瞻性全流水线评估是这项工作的下一阶段。
英文摘要
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
发表机构
- Harvard Medical School(哈佛医学院)
- Wayne State University School of Medicine(韦恩州立大学医学院)
- Women & Infants Hospital of Rhode Island(罗德岛妇女与婴儿医院)
- Warren Alpert Medical School of Brown University(布朗大学沃伦·阿尔珀特医学院)
- Duke University School of Medicine(杜克大学医学院)
机构由 AI 辅助整理,请以论文原文为准。