arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PET/CT放射报告的错误检测:领域特定模型与大语言模型

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

Hermione Warr, Harry Anthony, Lilli J Freischem, Yasin Ibrahim, Daniel R McGowan, Konstantinos Kamnitsas

arXiv 2608.30021首次发表:更新:

发表机构

University of Oxford; Oxford University Hospitals NHS; Imperial College London; University of Birmingham(牛津大学; 牛津大学医院NHS; 伦敦帝国理工学院; 伯明翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对PET/CT报告错误检测,对比了领域特定模型与大语言模型,发现15M参数的领域特定BERT模型性能优于最强提示式LLM,任务适配的Llama-3.3-70B可缩小差距但计算需求更高,证明领域特定训练更关键。

AI 中文摘要

放射报告中的错误会对患者治疗产生不利影响,但自动化报告质量保证仍具挑战性,因为错误往往较为细微,需领域专业知识才能检测。尽管近期有研究提出用大语言模型(LLMs)进行放射报告验证,但其在胸部X光数据集之外检测临床有意义错误的能力仍未得到充分探索。为此,本文首次对PET/CT报告错误检测的语言模型开展系统评估,对比了紧凑领域特定模型与SOTA开放权重LLMs。我们收集了23位放射科医生10年间的30633份肿瘤FDG PET/CT报告,训练了领域特定BERT模型以检测临床导向的合成报告错误,并在11500份保留的基准报告上与零/少样本Qwen3-32B、Gemma-3-27B和Llama-3.3-70B一同评估。15M参数模型达到94.4%的平衡准确率,假阳性率为5.8%,而最强的提示式LLM仅为84.0%。对Llama-3.3-70B进行任务特定适配缩小了这一性能差距(达94.4%),但保留了显著更高的计算需求。研究结果表明,对于PET/CT报告错误检测,领域特定训练比模型规模更重要,支持紧凑模型作为自动化放射报告质量保证的准确且计算高效的方法。

英文摘要

Errors in radiology reports can adversely affect patient treatment, yet automated report quality assurance remains challenging because errors are often subtle and require domain expertise to detect. Although large language models (LLMs) have recently been proposed for radiology report verification, their ability to detect clinically meaningful errors beyond chest X-ray datasets remains under-explored. To this end, we present the first systematic evaluation of language models for PET/CT report error detection, comparing compact domain-specific models with SOTA open-weight LLMs. We collected 30,633 oncology FDG PET/CT reports from 23 radiologists over 10 years. We trained domain-specific BERT models to detect clinically motivated synthetic reporting errors and evaluated alongside zero-/few-shot Qwen3-32B, Gemma-3-27B and Llama-3.3-70B on a held-out benchmark of 11,500 reports. A 15M-parameter model achieved 94.4% balanced accuracy with a 5.8% false-positive rate, compared with 84.0% for the strongest prompted LLM. Task-specific adaptation of Llama-3.3-70B closed this performance gap (94.4%) but retained substantially greater computational requirements. Our results suggest that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑