arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体流水线:从纵向结构化电子健康记录生成基于来源的合成临床笔记

A Multi-Agent Pipeline for Source-Grounded Synthetic Note Generation from Longitudinal Structured EHR

Nina Fatehi, Reihaneh Hassanzadeh, Meysam Ghaffari, Animesh Agarwal, Carlos Morato

arXiv 2609.22164首次发表:更新:

发表机构

Optum AI, UnitedHealth Group(Optum AI,联合健康集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MedNotes多智能体流水线,将纵向结构化EHR转化为基于来源的临床笔记,通过生成-评估-路由闭环实现质量控制,在EHRSHOT上达到91.4%通过率,并提升下游预测性能。

AI 中文摘要

结构化电子健康记录(EHR)数据丰富但稀疏、编码化,难以直接用于以笔记为中心的临床建模。我们提出了MedNotes,一个多智能体合成数据生成流水线,在明确的质量控制下将纵向结构化EHR转换为基于来源的临床笔记表示。MedNotes将结构化数据到文本的合成视为一个闭环智能体过程:生成器提出笔记,评估智能体诊断事实性、覆盖度、结构性和幻觉相关的失败,路由器接受、修订或拒绝草稿。在1,485个EHRSHOT就诊记录上,MedNotes实现了91.4%的通过率,平均事实准确性为0.980,完整性为99.1%,结构保真度为0.761,每次就诊的关键幻觉为0.028。迭代细化将接受率从69.4%提高到91.4%。生成的合成语料库在与有限真实数据结合时,改善了下游CPT预测和段落级章节预测。

英文摘要

Structured EHR is abundant but sparse, coded, and difficult to use directly for note-centric clinical modeling. We present MedNotes, a multi-agent synthetic data generation pipeline that converts longitudinal structured EHR into source-grounded clinical note representations under explicit quality control. MedNotes treats structured-data-to-text synthesis as a closed-loop agentic process: a generator proposes a note, evaluator agents diagnose factual, coverage, structural, and hallucination-related failures, and a router accepts, revises, or rejects the draft. On 1,485 EHRSHOT encounters, MedNotes achieves a 91.4% pass rate, with mean factual accuracy of 0.980, completeness of 99.1%, structural fidelity of 0.761, and 0.028 critical hallucinations per encounter. Iterative refinement improves acceptance from 69.4% to 91.4%. The resulting synthetic corpus improves downstream CPT prediction and paragraph-level section prediction when combined with limited real data.

CommentsAccepted at the FMSD Workshop, ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑