arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进超声报告的O-RADS风险分层:混合式与端到端大语言模型推理策略的对比评估

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

Xiaotong Tan, Chunli Qiu, Xin Liu, Qing Huang, Guangli Zhou, Bo Gao, Xiaoyan Song, Shuyan Wang, Xiuqin Wang, Wufeng Xue, Ruobing Huang, Dong Ni, Guowei Tao, Jun Cheng

arXiv 2608.23061首次发表:更新:

AI 中文总结

该研究对比不同LLM推理策略,发现将特征提取与规则分类解耦的混合架构,使用Gemini 3.6 Flash时O-RADS分类准确率达99.2%,优于端到端策略及原始临床报告,可实现准确可靠的自动化临床决策。

AI 中文摘要

背景:用大语言模型(LLM)实现基于临床指南的决策自动化仍面临可靠性、幻觉及可解释性有限的挑战。我们对比了LLM及推理策略,用于从自由文本盆腔超声报告中自动进行卵巢-附件报告与数据系统(O-RADS)分类。方法:本回顾性研究纳入接受盆腔超声检查的连续卵巢肿块患者,测试8种LLM的3种推理策略:隐式知识端到端、规则引导端到端,以及将特征提取与基于规则的分类解耦的基于特征的混合架构。参考标准为专家共识确定的O-RADS分类。结果:共评估310名女性的390个卵巢肿块。使用Gemini 3.6 Flash的基于特征的混合架构表现最佳,准确率达99.2%(390个中的387个),与参考标准几乎完全一致(加权kappa=1.00;95% CI:0.99-1.00)。其性能优于原始临床报告(准确率87.7%,390个中的342个;加权kappa=0.94;95% CI:0.91-0.96)及端到端LLM策略(准确率范围65.6%,390个中的256个,至95.9%,390个中的374个)。在结构化特征提取方面,Gemini 3.6 Flash的整体准确率高于Claude Fable 5(98.9% vs 97.8%;P<0.001)。该混合架构减少了分类错误,缓解了原始报告中观察到的过度分期倾向。结论:将临床特征提取与确定性指南执行分离的基于特征的混合LLM架构,可实现高度准确、可靠且可解释的自动化O-RADS分类,为标准化、基于指南的临床决策提供了有前景的方法。

英文摘要

Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P < 0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

CommentsMain manuscript: 20 pages, 5 figures, and 2 tables; supplemental material: 11 pages, 1 figure, and 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑