arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05425cs.CVcs.CLcs.LG

未提及的检查发现改变了强化学习在胸部X光报告检查中的表现评估

Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

Ali Vosoughi, Akhil Kasturi, Chenliang Xu, Axel Wismueller

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过强化学习训练视觉-语言模型生成胸部X光报告检查清单,发现未提及发现的存在显著影响检查模型性能,提示评估方法需考虑此因素。

中文摘要 AI 辅助

放射学报告的自动化检查可能依赖于AI生成的清单,而这些清单会遗漏未提及的发现。我们使用强化学习训练了一个视觉-语言模型,在不查看待测句子的情况下,从胸部X光片填写一个包含12项发现的清单;一个独立的检查模型根据清单对句子进行判断。在留出患者上,一个基于规则的检查和一位独立的医学检查者(均未在训练中使用)测量的判别增益(Youden指数)分别为12.6%和11.8%;只有基于规则的检查达到了预设的误报标准。切换到训练格式(该格式固定了发现顺序并将未提及的发现视为缺失)后,训练检查者测量的增益提高,而独立检查者的增益降低,这一预设比较产生了6.2%(95%区间2.0%至10.5%)的差异,并且在留出患者上的事后分析中为7.7%。在8个检查模型中,对未提及发现的标签一致否定陈述的接受率范围从1.0%到97.0%。标签来源于报告,而非放射科医生裁定。

英文摘要

Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.

发表机构

  • University of Rochester(罗切斯特大学)
  • University of Rochester Medical Center(罗切斯特大学医学中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑