未提及的检查发现改变了强化学习在胸部X光报告检查中的表现评估
Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking
浏览论文内容
中文总结 AI 辅助
本研究通过强化学习训练视觉-语言模型生成胸部X光报告检查清单,发现未提及发现的存在显著影响检查模型性能,提示评估方法需考虑此因素。
中文摘要 AI 辅助
放射学报告的自动化检查可能依赖于AI生成的清单,而这些清单会遗漏未提及的发现。我们使用强化学习训练了一个视觉-语言模型,在不查看待测句子的情况下,从胸部X光片填写一个包含12项发现的清单;一个独立的检查模型根据清单对句子进行判断。在留出患者上,一个基于规则的检查和一位独立的医学检查者(均未在训练中使用)测量的判别增益(Youden指数)分别为12.6%和11.8%;只有基于规则的检查达到了预设的误报标准。切换到训练格式(该格式固定了发现顺序并将未提及的发现视为缺失)后,训练检查者测量的增益提高,而独立检查者的增益降低,这一预设比较产生了6.2%(95%区间2.0%至10.5%)的差异,并且在留出患者上的事后分析中为7.7%。在8个检查模型中,对未提及发现的标签一致否定陈述的接受率范围从1.0%到97.0%。标签来源于报告,而非放射科医生裁定。
英文摘要
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.
发表机构
- University of Rochester(罗切斯特大学)
- University of Rochester Medical Center(罗切斯特大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。