ASTAR:从大规模临床自由文本语料库自动生成标准化放射学报告模板
ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora
- Tsinghua University(清华大学)
- Sichuan University(四川大学)
- University of California San Diego(加利福尼亚大学圣迭戈分校)
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对放射学报告模板构建的人工瓶颈,提出基于LLM的ASTAR框架,从大规模临床自由文本语料库自动生成标准化模板,实验显示其性能优于专家构建的模板,大幅缩短模板开发时间。
AI中文摘要:
结构化报告将自由文本放射学叙述转换为可查询的数据键,便于队列组装、纵向跟踪以及为医学AI生成训练标签。当前主流范式采用两阶段流程:(1)构建报告模板;(2)提取信息以填充模板。尽管提取阶段已受益于大型语言模型(LLMs)的进展,但模板构建仍为人工瓶颈,依赖耗时费力的专家共识,这种共识具有静态性、难以规模化的特点,且可能无法捕捉现实世界报告的多样性。我们针对该局限性提出了基于LLM的框架ASTAR,用于从大规模临床自由文本语料库自动生成标准化放射学报告模板。对来自多个中心的4215份胎儿脑部MRI报告开展的大量实验表明,ASTAR生成的模板在模板覆盖率、信息保真度、诊断保真度以及专家评估可用性方面均优于两个专家精心构建的模板,将模板开发从数周的委员会审议缩短至数小时的自动处理。代码:this https URL
英文摘要:
Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constructing a reporting template, (2) extracting information to populate it. While the extraction stage has benefited from advances in large language models (LLMs), template construction remains a manual bottleneck relying on labor-intensive expert consensus that is static, difficult to scale, and may fail to capture real-world reporting diversity. We address this limitation with \textbf{\texttt{ASTAR}}, an LLM-based framework for Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora. Extensive experiments on 4,215 fetal brain MRI reports from multiple centers demonstrate that the \textbf{\texttt{ASTAR}}-induced template surpasses two expert-curated templates across template coverage, information fidelity, diagnostic fidelity, and expert-rated usability, reducing template development from weeks of committee deliberation to hours of automated processing. Code: https://github.com/birthlab/ASTAR