arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型从影像申请文本中提取的文本嵌入实现胸部CT协议自动选择

Automated Chest CT Protocol Selection via Large Language Model Derived Text Embeddings from Imaging Request Text

Zahra Hosseini, Mahan Pouromidi, Farzad Khalvati, Patrik Rogalla

arXiv 2609.07986首次发表:更新:

发表机构

University Health Network; Toronto General Hospital; The Hospital for Sick Children; University of Toronto(大学健康网络; 多伦多综合医院; 病童医院; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用微调LLM(LLaMA-3.1-70B)嵌入临床文本,结合逻辑回归分类器自动选择胸部CT协议,在285,123份申请上达到79%准确率,与放射科医生表现相当,可减少选择变异性。

AI 中文摘要

目的:准确的CT协议选择对于诊断质量和患者安全至关重要,然而当前流程是人工操作,耗时且容易出现不一致。以往使用关键词或词袋的机器学习方法缺乏上下文理解能力,在罕见协议上表现不佳。我们提出了一种决策支持系统,利用大语言模型(LLM)特征从自由文本临床指征中推荐协议,捕捉临床细微差别和措辞变化,以实现更一致、更高效的选择。方法:在这项经REB批准的回顾性研究中,来自一家大型学术医疗中心(2017-2024年)的285,123份胸部CT影像申请被分为训练集(228,099份,80%)和留出测试集(57,024份,20%)。每份申请包括手术名称、临床指征、HIS评论和所选协议。临床文本使用微调后的LLM(Meta的LLaMA-3.1-70B)进行嵌入;这些特征输入逻辑回归分类器,预测18种协议标签(例如,PE、LDCT)。结果:该流程在18种CT协议上实现了加权精确率0.84、加权F1分数0.81和总体准确率79%。在300例具有专家共识的独立病例中,LLM的总体准确率达到80%,而放射科医生为83%,两者无显著差异(p = 0.263)。大多数类别的性能相当,LLM在某些具有挑战性的类别上超过放射科医生,熵分析表明协议使用更均衡,提示变异性降低。结论:基于LLM的推荐系统可以利用大型自然文本语料库中的通用知识,从自由文本影像申请中准确分配胸部CT协议,并可能作为需要语言理解的协议推荐工具的可行基础。

英文摘要

Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current process is manual, time-consuming, and prone to inconsistencies. Prior Machine Learning methods using keywords or bag-of-words lack contextual understanding and perform poorly on rare protocols. We propose a decision support system using large language model (LLM) features to recommend protocols from free-text clinical indications, capturing clinical nuance and phrasing variation for more consistent, efficient selection. Methods: In this REB-approved retrospective study, 285,123 chest CT imaging requests from a large academic medical center (2017-2024) were split into training (228,099, 80%) and held-out test (57,024, 20%) sets. Each request included procedure names, clinical indication, HIS comments, and the selected protocol. Clinical text was embedded using a fine-tuned LLM, Meta's LLaMA-3.1-70B; these features input a logistic regression classifier predicting 18 protocol labels (e.g., PE, LDCT). Results: The pipeline achieved a weighted precision of 0.84, weighted F1-score of 0.81, and overall accuracy of 79% across 18 CT protocols. On 300 independent cases with expert consensus, the LLM reached an overall accuracy of 80% versus 83% for radiologists, with no significant difference (p = 0.263). Performance was comparable across most classes, with the LLM exceeding radiologists for some challenging categories, and entropy analyses indicated more balanced protocol use, suggesting reduced variability. Conclusion: An LLM-based recommendation system can leverage general knowledge from a large natural-text corpus to accurately assign chest CT protocols from free-text imaging requests, and may serve as a viable foundation for protocol recommendation tools where inputs require language understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑