arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

乳腺X线摄影和乳腺超声报告中临床发现的提取:专家与人工智能的比较

Extraction of clinical findings from mammography and breast ultrasound reports: a comparison between specialists and Artificial Intelligence

Lorenzo Farias, Hanna Reckziegel, Daniela Duarte da Silva Bagatini, Daniel Schulz, Gabriela de Andrade Monteiro, Letícia Zanatta, Ana Laura Brill Thum, Priscila Schmidt Lora, Débora Oliveira da Silva, Ana Paula Wernz da Cunha Müller, Cristiane Drebes Pedron

arXiv 2609.31974首次发表:更新:

AI 中文总结

本研究比较LLM与专家在巴西乳腺影像报告中提取临床发现的性能,Gemini 2.5 Flash模型通过少样本提示工程实现宏F1 0.91,优于人工提取,并可作为辅助验证工具。

AI 中文摘要

乳腺癌是巴西女性癌症相关死亡的主要原因,从申请到发布乳腺X线摄影报告之间的时间直接影响筛查的依从性,因此处理这些报告的敏捷性是早期诊断的关键因素。在此背景下,本研究比较了大语言模型(LLM)与由健康研究人员组成的人工提取团队在识别巴西葡萄牙语撰写的乳腺X线摄影和乳腺超声报告中临床发现方面的性能。通过使用少样本策略的提示工程应用了命名实体识别(NER),采用了从四个候选模型的初步探索性测试中选出的Gemini 2.5 Flash模型。Gemini 2.5 Flash模型表现出最佳性能,实现了宏F1为0.91和微F1为0.98。主观验证中,健康研究人员使用李克特量表评估了29份不同格式的检查,得出了一致性指数为93.1%。该模型在整体宏F1上优于人工提取(0.91对0.72),并且在四份报告中,模型正确识别了人工提取期间被遗漏或错误记录的信息,展示了其作为专家辅助验证工具的潜力。结果证实了假设,即通过提示工程指导的LLM可以实现与健康专业人员人工提取相当或更优的性能。

英文摘要

Breast cancer is the leading cause of cancer-related death among women in Brazil, and the time between the request and the release of mammography reports directly influences adherence to screening, making the agility in processing these reports a critical factor for early diagnosis. In this context, this study compares the performance of a Large Language Model (LLM) with manual extraction performed by a team of health researchers in identifying clinical findings from mammography and breast ultrasound reports written in Brazilian Portuguese. Named Entity Recognition (NER) was applied through Prompt Engineering using a few-shot strategy, employing the Gemini 2.5 Flash model, selected from preliminary exploratory tests with four candidate models. The Gemini 2.5 Flash model demonstrated the best performance, achieving a Macro F1 of 0.91 and a Micro F1 of 0.98. The subjective validation, in which 29 exams of different formats were evaluated by health researchers using a Likert scale, yielded an agreement index of 93.1%. The model outperformed human extraction in overall Macro F1 (0.91 vs. 0.72), as in four reports the model correctly identified information that had been omitted or incorrectly recorded during manual extraction, demonstrating its potential as a complementary verification tool alongside specialists. The results confirm the hypothesis that LLMs, when instructed through Prompt Engineering, can achieve performance comparable to or superior to manual extraction by health professionals.

Comments16 pages, 10 figures, 3 tables. Submitted to Artificial Intelligence in Medicine

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑