发表机构
Department of Anatomical Pathology, Singapore General Hospital; Duke–NUS Medical School; University of Illinois Urbana-Champaign; Independent Researcher(新加坡中央医院解剖病理科; 新加坡国立大学Duke-NUS医学院; 伊利诺伊大学厄巴纳-香槟分校; 独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对胃活检报告中幽门螺杆菌证据提取难题,采用Nimblemind多智能体系统,经试点评估其总体准确率达98.61%,虽与其他比较器性能相似,但实现了工作流程整合和可追溯性,还能大幅减少审查时间。
AI 中文摘要
新加坡数据显示约31%人口有幽门螺杆菌感染迹象,其持续感染与多种疾病相关,根除是预防胃癌关键。但支持幽门螺杆菌阳性及相关胃炎的证据分布在不同编码和自由文本字段,限制关键词搜索且人工审查难扩展。我们对Nimblemind多智能体系统(nMAS)进行回顾性试点评估,用54份匿名胃活检病理报告,评估四个临床二元字段。nMAS总体准确率达98.61%,与另一个比较器性能相似,但nMAS能保持统一报告级输出及支持源句子,展示了工作流程整合和可追溯性。在假设情景下,证据关联验证可大幅减少审查时间,更大规模多机构研究应评估证据跨度正确性、临床医生验证时间和通用性。
英文摘要
Clinical feature extraction from pathology reports is challenging because relevant evidence may be distributed across coded and narrative fields and depend on specimen attribution, negation, ancillary findings, and diagnostic context. We retrospectively evaluated the NimbleMind Multi-Agent System (nMAS), a configurable workflow that separates clinician-defined field specifications from extraction models and returns report-level predictions with source-linked evidence. The study included 54 dummy gastric biopsy pathology reports from Singapore and four binary target fields, yielding 216 feature-case decisions. nMAS correctly classified 213 of 216 decisions (98.61\%), and all evidence spans associated with correct predictions occurred verbatim in the corresponding source reports. All three errors occurred in the two context-dependent \textit{H. pylori}-related fields requiring negation handling or diagnostic attribution. A single-model UMA-style comparator produced the similar label-level performance and error pattern. These findings do not demonstrate predictive superiority for the multi-agent architecture.Rather, the contribution of nMAS lies in workflow integration and traceability through configurable field specifications, complexity-based routing, report-level aggregation, and source-text validation within a clinician-reviewable workflow. Larger multi-institutional studies should assess generalizability, semantic evidence quality, adaptation effort, and clinician verification time.