发表机构
Nimblemind; School of Interactive Computing, Georgia Institute of Technology; NYU Tandon School of Engineering, New York University; Florida International University; Duke-NUS Medical School; Singapore General Hospital; University of Illinois Urbana-Champaign; OncoLens(Nimblemind; 佐治亚理工学院交互计算学院; 纽约大学坦登工程学院; 佛罗里达国际大学; 杜克-新加坡国立大学医学院; 新加坡中央医院; 伊利诺伊大学厄巴纳-香槟分校; OncoLens)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对肿瘤学文档的结构化提取需求,提出nMAS多智能体工作流,在230份肿瘤学文档上的F1值达85.0%,优于对照模型,验证了其将碎片化肿瘤学文档转为结构化数据的可行性。
AI 中文摘要
临床相关的肿瘤学信息分布在异构的纵向文档中,这造成了巨大的抽象负担,需要准确关联标本、肿瘤、生物标志物和时间点,而手动癌症登记抽象每个病例可能需要27.2分钟,凸显了对可扩展方法的需求,该方法需在保留临床背景的同时将文档转换为结构化数据。我们评估了Nimblemind多智能体系统(nMAS),这是一种可配置的肿瘤学信息提取工作流,可从碎片化的肿瘤学文档中提取临床相关的结构化字段。提取任务采用了临床医生制定的包含328个属性的模式,涵盖报告元数据、诊断、分期和癌症类型特定信息。nMAS将临床医生定义的字段规范与模型执行分离,并结合了感知复杂性的提取、报告级别的整合以及基于源的验证。回顾性评估纳入了来自40名患者的230份去标识化肿瘤学文档,以及418个经临床医生审核的包含1126个非空参考值的文档-字段对。评估聚焦于临床医生确定的源文档中存在的字段,而非详尽注释全部328个模式字段。nMAS实现了秩加权值级别的精确率为82.6%、召回率为87.5%、F1值为85.0%,相比之下,独立实现的UMA式MiniMax M2.5对照模型的F1值为66.4%。这些发现支持使用可配置的、基于源的提取工作流将碎片化肿瘤学文档转换为可重复使用的结构化数据的可行性。
英文摘要
Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual cancer-registry abstraction can require 27.2 minutes per case, highlighting the need for scalable methods that preserve clinical context while converting documentation into structured data. We evaluate an oncology information-extraction workflow in which OncoLens supplies multi-source, oncology-aware document selection, aggregation, and normalization from integrated EHRs, while the NimbleMind Multi-Agent System (nMAS) is a configurable oncology information-extraction workflow that extracts clinically relevant structured fields from fragmented oncology documentation. The extraction task uses a clinician-informed schema of 328 attributes spanning report metadata, diagnosis, staging, and cancer-type-specific information. nMAS separates clinician-defined field specifications from model execution and combines complexity-aware extraction, report-level consolidation, and source-grounded validation. The retrospective evaluation included 230 de-identified oncology documents from 40 patients and 418 clinician-reviewed document-field pairs containing 1,126 non-empty reference values. Evaluation focused on fields identified by clinicians as present in the source documents rather than exhaustively annotating all 328 schema fields. nMAS achieved a rank-weighted value-level precision of 82.6%, recall of 87.5%, and F1 of 85.0%, compared with an F1 of 66.4% for an independently implemented UMA-style MiniMax M2.5 comparator. These findings support the feasibility of using a configurable, source-grounded extraction workflow to convert fragmented oncology documentation into reusable structured data.