发表机构
Emory University School of Medicine; Emory University(埃默里大学医学院; 埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较微调开放权重模型Gemma-3-12B与GPT-4o在颅内出血严重程度提取中的表现,发现蒸馏真实报告数据而非合成数据是匹配托管模型的关键,实现了私有低成本的本地替代方案。
AI 中文摘要
将自由文本放射学报告转换为结构化标签,有助于队列构建、质量保证和临床影像模型的监控,但最强的标签提取器是托管的专有模型,其使用引发了隐私、成本和可重复性问题。我们研究了一个微调后的开放权重模型(Gemma-3-12B)能否在非对比头部CT报告的多标签颅内出血(ICH)严重程度提取中匹配GPT-4o,以及哪些因素至关重要。采用2x2设计,我们将两种适应策略(判别式分类头CH;生成式指令微调IFT)与两种训练数据来源(从真实GPT-4o标记报告中蒸馏;由GPT-4o基于真实样本生成的合成报告)交叉组合,跨越五种训练规模,在100份专家裁决报告上对照GPT-4o和未微调的开放权重基础模型进行基准测试。蒸馏指令微调模型(DIFT)匹配了GPT-4o(宏F1 0.845对0.850;p=1.000),并超过基础模型0.178。决定性因素是训练数据来源,而非微调方法:两种合成数据模型在任何训练规模下均未超过未微调的开放权重基础模型,并在所有严重程度类别上表现逊于蒸馏模型。微调和推理均适配于单个24GB消费级GPU的内存范围。对于狭窄、高价值的临床标签提取任务,蒸馏真实报告而非生成合成报告,是缩小与托管模型差距的关键,从而提供一种私有、低成本、版本稳定的本地替代方案。
英文摘要
Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.