arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双重建议与单一建议AI支持对住院医师放射影像解读的影响:随机多读者研究

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng, Chunlong Zhang, Yuting Zhu, Kaixiao Chen, Xiao Liang, Chen Yang, Yeyuan Chen, Hao Chen, Fuqing Zhou

arXiv 2610.09589首次发表:更新:

发表机构

Nanchang University; Jiangxi Province Medical Imaging Research Institute; Hong Kong University of Science and Technology(南昌大学; 江西省医学影像研究所; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过随机多读者试验发现,双重AI建议相比单一建议能显著提升放射科住院医师的影像解读准确率,尤其在AI建议错误时,但对非放射科医师无显著影响。

AI 中文摘要

目的:比较双重建议与单一建议AI支持对住院医师放射影像解读的影响,特别是在共享AI建议错误的情况下。材料与方法:这项前瞻性、多中心、随机三臂读者研究于2026年7月至9月在中国三家医院进行(ChiCTR2600129243)。经过专业分层后,132名临床经验不足3年的住院医师按1:1:1随机分配至仅使用GPT-5.4(A组)、GPT-5.4加Kimi-K2.6(B组)或GPT-5.4加Gemini-3.6 Flash(C组);其中123名被纳入分析。参与者在AI支持前后解读了60张X光片。主要结果是准确率变化。使用Welch ANOVA和Holm校正t检验比较支持条件;使用HC3线性模型评估专业交互作用。结果:在123名住院医师中(平均年龄24.1岁±1.4;65名女性),放射科住院医师在双重建议支持下的准确率提升大于单一建议支持(B-A,6.69个百分点[95% CI,0.97-12.40];C-A,7.87个百分点[95% CI,1.64-14.11];Holm校正P=.030),而非放射科住院医师的准确率变化无显著差异(P=.20)。当GPT-5.4错误时,放射科住院医师在双重建议支持下的AI辅助准确率高于单一建议支持(40.1%和40.4%对20.0%),非放射科住院医师也如此(31.3%和31.0%对12.1%)(所有Holm校正P<.001)。双重建议效应因专业而异(交互差异,10.44个百分点;95% CI,4.36-16.52;P<.001)。结论:双重建议支持可能减轻错误AI建议的影响,在放射科住院医师中观察到更大的准确率提升,但在非放射科住院医师中未观察到。

英文摘要

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑