基础智能体与智能型深度研究结合:基于证据的临床代码预测
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
浏览论文内容
中文总结 AI 辅助
本文提出ICD-Deepresearch工作流,结合EHR基础模型、语言基础模型与医学搜索工具,在MIMIC-III、MIMIC-IV数据集上实现ICD编码预测,其检索证据的医生有用性评分优于对比系统。
中文摘要 AI 辅助
下一次就诊的ICD编码预测任务是,利用此前已有的纵向记录,预测未来就诊时将记录哪些标准化诊断编码,该任务具有前瞻性且为多标签类型:目标就诊记录尚未存在,可能有多个编码正确。结构化电子健康记录(EHR)基础模型可捕捉疾病复发与时间进展,而语言基础模型能生成灵活的诊断假设。本文提出ICD-Deepresearch,这一深度研究工作流将这些预测性基础模型与医学搜索工具、ICD编码词典相结合。由于无任何数据源可揭示未来的编码集合,该研究在固定的Top-K预算下,通过关联患者证据、外部临床关系及编码的精确语义来评估候选转换。候选生成环节使用SparseEHR生成EHR先验,初始化两轮受限的研究扩展;独立的GPT-5直接预测补充候选。最终选择环节验证、去重并联合排序两条路径,之后单独模块生成理由但不改变预测结果。ICD-Deepresearch在MIMIC-III数据集上实现患者平均精确率/召回率为24.60%/35.09%,在MIMIC-IV数据集上为25.14%/48.32%;医生对其检索文档的有用性评分分别为51%和68%,而独立GPT-5网络搜索的对应评分仅为22%和39%,Medical Deep Research的对应评分为32%和41%。因此ICD-Deepresearch相较于已注册的本地对比系统有所提升,且检索到的证据获得的医生评分有用性高于独立研究系统。
英文摘要
Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems