AI 中文总结
本研究开发了一款本地部署的多智能体AI系统,可实现放射科报告结构化与质量保证,经独立放射科医师评估,该系统表现良好,或有助于放射科报告标准化。
AI 中文摘要
目的:开发并评估一款本地部署的多智能体AI系统,用于放射科报告的结构化与质量保证(QA)。材料与方法:本回顾性研究纳入了2023年至2024年间15名持证放射科医师口述的638份胸部、腹部及盆腔CT检查放射科报告。研究人员开发了一套多智能体AI流程,用于执行报告结构化与质量保证任务。该系统利用正则表达式规则和本地大语言模型,将报告在句子层面结构化至标准化解剖学章节;同时可检测“发现”与“印象”章节间或章节内的不匹配、性别-解剖学冲突,以及关键发现的未记录沟通情况。两名持证放射科医师独立评估了45份报告的子集。结果:该多智能体系统将所有报告的“发现”章节(共22270个句子)结构化至预定义的解剖学格式,同时保留了报告的原始内容;该系统标记了90份(占比14.1%)报告,最常见的标记原因为章节不匹配(80份,占比12.5%)。在放射科医师评估中,两名评审医师一致认为31份(占比69%)报告被正确结构化,2份(占比4%)被错误结构化,剩余12份(占比27%)的评估意见存在分歧;两名医师均认为未遗漏任何临床重要信息,也未引入虚构内容。在接受评估的报告中,整体QA表现被评为“优秀”或“良好”的占84%,其余报告被评为“一般”。结论:一款本地部署的多智能体AI系统将放射科报告结构化与质量保证整合至单一工作流程中,该系统在放射科医师评估中展现出良好性能,此类系统或可支持放射科实践中的报告标准化与质量保证工作。
英文摘要
Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.
Comments14 pages, 2 figures, 4 tables