发表机构
Nanchang University; Jiangxi Province Medical Imaging Research Institute; Hong Kong University of Science and Technology(南昌大学; 江西省医学影像研究所; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过随机多读者试验发现,双重AI建议相比单一建议能显著提升放射科住院医师的影像解读准确率,尤其在AI建议错误时,但对非放射科医师无显著影响。
AI 中文摘要
目的:比较双重建议与单一建议AI支持对住院医师放射影像解读的影响,特别是在共享AI建议错误的情况下。材料与方法:这项前瞻性、多中心、随机三臂读者研究于2026年7月至9月在中国三家医院进行(ChiCTR2600129243)。经过专业分层后,132名临床经验不足3年的住院医师按1:1:1随机分配至仅使用GPT-5.4(A组)、GPT-5.4加Kimi-K2.6(B组)或GPT-5.4加Gemini-3.6 Flash(C组);其中123名被纳入分析。参与者在AI支持前后解读了60张X光片。主要结果是准确率变化。使用Welch ANOVA和Holm校正t检验比较支持条件;使用HC3线性模型评估专业交互作用。结果:在123名住院医师中(平均年龄24.1岁±1.4;65名女性),放射科住院医师在双重建议支持下的准确率提升大于单一建议支持(B-A,6.69个百分点[95% CI,0.97-12.40];C-A,7.87个百分点[95% CI,1.64-14.11];Holm校正P=.030),而非放射科住院医师的准确率变化无显著差异(P=.20)。当GPT-5.4错误时,放射科住院医师在双重建议支持下的AI辅助准确率高于单一建议支持(40.1%和40.4%对20.0%),非放射科住院医师也如此(31.3%和31.0%对12.1%)(所有Holm校正P<.001)。双重建议效应因专业而异(交互差异,10.44个百分点;95% CI,4.36-16.52;P<.001)。结论:双重建议支持可能减轻错误AI建议的影响,在放射科住院医师中观察到更大的准确率提升,但在非放射科住院医师中未观察到。
英文摘要
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.