发表机构
Charité – Berlin University of Medicine; Technical University of Munich; University of Salerno; Aristotle University of Thessaloniki; Berlin University of Applied Sciences and Technology; Berlin Institute of Health(柏林夏里特医学院; 慕尼黑工业大学; 萨莱诺大学; 亚里士多德大学; 柏林应用技术大学; 柏林健康研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发并验证了一种基于开放权重大语言模型gpt-oss-120B的流水线,能在无需人工监督的情况下,以高保真度将大规模多模态放射学报告档案自动结构化,但多区域报告仍是主要错误来源。
AI 中文摘要
目的:开发并评估一种开放权重大语言模型(LLM)流水线,该流水线可在无需人工监督的情况下,将整个自由文本放射学报告档案转换为结构化报告。材料与方法:在这项回顾性研究中,在一个中心开发了包含150个分层组织模板的流水线,并在另一个中心对2010年至2025年的报告进行了测试。开放权重模型gpt-oss-120B通过三个约束解码步骤选择模板,并在单个本地图形处理单元(GPU)上填充模板。模板选择依据专家标签对914份随机抽样的五种模态报告进行评分,结构化质量则由五名住院医师对920份X线摄影和CT报告逐字段修正后进行评分。随后,该流水线处理了第二个中心的完整档案。比例以Wilson 95%置信区间(CI)报告。结果:为74.4%的报告(680/914;95% CI:71.5%,77.1%)选择了最佳模板集,为82.3%(752/914;95% CI:79.7%,84.6%)选择了合适模板集,其中单区域报告为87.7%,多区域报告为54.1%。输出与修正后参考之间的宏观语义文本相似度在X线摄影中为0.95,在CT中为0.97;住院医师对24,638个字段中的88.7%未作修改,不支持的内容在1.0%和1.5%的报告中被标记。在2,186,982份档案报告中,96.5%获得了结构化输出,即2,401,544份结构化报告,处理速度为每小时1,258份报告(使用单个GPU)。结论:开放权重LLM流水线在无需人工监督的情况下,以高内容保真度结构化了一个完整的多模态报告档案。多区域报告仍是模板错误的主要来源。
英文摘要
Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline with 150 hierarchically organized templates was developed at one center and tested at a second center on reports from 2010 to 2025. The open-weight model gpt-oss-120B selects the template in three constrained-decoding steps and fills it on one local graphics processing unit. Template selection was scored against expert labels on 914 randomly sampled reports of five modalities, structuring quality on 920 radiography and CT reports corrected field by field by five residents. The pipeline then processed the complete archive of the second center. Proportions are reported with Wilson 95% confidence intervals (CIs). Results: An optimal template set was selected for 74.4% of reports (680 of 914; 95% CI: 71.5%, 77.1%) and an appropriate set for 82.3% (752 of 914; 95% CI: 79.7%, 84.6%), 87.7% for single-region and 54.1% for multi-region reports. Macro semantic textual similarity between output and corrected reference was 0.95 for radiography and 0.97 for CT, residents left 88.7% of 24,638 fields unchanged, and unsupported content was flagged in 1.0% and 1.5% of reports. Of 2,186,982 archive reports, 96.5% received structured output, 2,401,544 structured reports, at 1,258 reports per hour on one graphics processing unit. Conclusion: An open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity. Multi-region reports remained the main source of template errors.
Comments27 pages, 9 figures