AI 中文总结
该研究对比不同LLM推理策略,发现将特征提取与规则分类解耦的混合架构,使用Gemini 3.6 Flash时O-RADS分类准确率达99.2%,优于端到端策略及原始临床报告,可实现准确可靠的自动化临床决策。
AI 中文摘要
背景:用大语言模型(LLM)实现基于临床指南的决策自动化仍面临可靠性、幻觉及可解释性有限的挑战。我们对比了LLM及推理策略,用于从自由文本盆腔超声报告中自动进行卵巢-附件报告与数据系统(O-RADS)分类。方法:本回顾性研究纳入接受盆腔超声检查的连续卵巢肿块患者,测试8种LLM的3种推理策略:隐式知识端到端、规则引导端到端,以及将特征提取与基于规则的分类解耦的基于特征的混合架构。参考标准为专家共识确定的O-RADS分类。结果:共评估310名女性的390个卵巢肿块。使用Gemini 3.6 Flash的基于特征的混合架构表现最佳,准确率达99.2%(390个中的387个),与参考标准几乎完全一致(加权kappa=1.00;95% CI:0.99-1.00)。其性能优于原始临床报告(准确率87.7%,390个中的342个;加权kappa=0.94;95% CI:0.91-0.96)及端到端LLM策略(准确率范围65.6%,390个中的256个,至95.9%,390个中的374个)。在结构化特征提取方面,Gemini 3.6 Flash的整体准确率高于Claude Fable 5(98.9% vs 97.8%;P<0.001)。该混合架构减少了分类错误,缓解了原始报告中观察到的过度分期倾向。结论:将临床特征提取与确定性指南执行分离的基于特征的混合LLM架构,可实现高度准确、可靠且可解释的自动化O-RADS分类,为标准化、基于指南的临床决策提供了有前景的方法。
英文摘要
Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P < 0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.
CommentsMain manuscript: 20 pages, 5 figures, and 2 tables; supplemental material: 11 pages, 1 figure, and 3 tables