AI 中文总结
针对临床试验系统文献综述传统方法的不足,提出含人在回路的两个多智能体系统,筛选MAS用多LLM智能体等提高准确性,提取MAS结合多种手段确保准确与可扩展,重现网络荟萃分析得出新临床结论。
AI 中文摘要
临床试验的系统文献综述推动监管决策,但传统的筛选和提取耗时、费力且易受研究选择偏差影响。我们提出了两个适用于系统文献综述的多智能体系统(MAS),且有人在回路。筛选MAS使用具有异构角色的多个大语言模型(LLM)智能体和多轮交叉评审,比单LLM基线均匀提高了准确性。提取MAS结合了标准化、迭代校正回路和基于检索的上下文控制以确保准确性和可扩展性。两个MAS专为支持人在回路而设计,这对临床决策至关重要。所提方法的新颖之处在于系统架构,而非任何单个基础工具,该系统可从基础工具的未来改进中自然受益。作为实际应用,MAS重现了已发表的网络荟萃分析。结果找回了原始研究中的所有试验,并识别出人工评审遗漏的其他合格试验,从而得出更新的临床结论。
英文摘要
Systematic literature review of clinical trials drives regulatory decision-making, but conventional screening and extraction are time-consuming, labor-intensive, and vulnerable to study selection bias. We propose two fit-to-purpose multi-agentic systems (MAS) for systematic literature review, with human-in-the-loop. The screening MAS uses multiple LLM agents with heterogeneous personas and multiround cross-review, and uniformly improves accuracy over a single-LLM baseline. The extraction MAS combines standardization, an iterative correction loop, and retrieval-based context control to ensure accuracy and scalability. Both MAS are specifically designed to support Human-In-The-Loop which is essential for clinical decisions. The novelty of the proposed approach lies in the system architecture rather than in any single foundation tools: the system can naturally benefit from future improvements in the underlying tools, for instance, stronger LLM agents, retrieval engines, image recognition methods, etc. As a real-world application, a published network meta-analysis is reproduced by the MAS. The result recovers all trials from the original study and identifies additional eligible trials missed by manual review, leading to updated clinical conclusions.