发表机构
University of Florida; NVIDIA; Lehigh University; Indiana University; Regenstrief Institute(佛罗里达大学; 英伟达; 理海大学; 印第安纳大学; 雷根斯特里夫研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出智能体生成式大语言模型GatorOnco,经大规模生物医学文本训练与领域适配,在结直肠癌治疗规划的临床评估中性能优于开源LLM,达到专家级水平,可缩小生成式AI在高风险诊疗规划领域的应用差距。
AI 中文摘要
精准肿瘤学中的治疗规划需要整合异质性患者信息与快速更新的临床指南,以确保符合指南的诊疗方案。尽管大语言模型(LLM)在诸多诊断任务中展现出应用前景,但其在高风险的治疗规划中的应用仍受限于复杂推理、对时效性临床指南的遵循以及安全性问题。本研究提出GatorOnco,一种用于结直肠癌(CRC)治疗规划的智能体大语言模型。GatorOnco基于总计2820亿个生物医学文本令牌开发,其中包括来自UF Health的1660亿个医疗系统规模的临床文本令牌。我们实现了一种领域适配方法,整合了预训练、模型合并、两阶段后训练方法以及基于智能体的强化学习。一种智能体检索增强生成(RAG)方法将时效性临床指南动态整合至推理过程中。在由5名UF Health肿瘤学家开展的盲法随机临床评估中,GatorOnco的表现显著优于开源LLM(P < 0.01),达到与UF Health肿瘤学家相当的专家级性能。与专家肿瘤学家相比,GatorOnco在可读性(4.46 vs. 4.19,P < 0.01)和完整性(3.91 vs. 3.52,P < 0.01)方面获得显著更高的评分,而在正确性(4.09 vs. 4.11,P = 0.921)、时效性(4.04 vs. 3.98,P = 0.478)和安全性(4.22 vs. 4.22,P = 0.999)方面表现出统计学上相当的性能。这些研究结果表明,将智能体推理与大规模领域适配相结合,可助力缩小生成式AI在高风险癌症治疗规划领域的应用差距。
英文摘要
Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnostic tasks, their adoption for high-stakes treatment planning is hindered by complex reasoning, adherence to timely clinical guidelines, and safety concerns. In this study, we present GatorOnco, an agentic LLM for colorectal cancer (CRC) treatment planning. GatorOnco is developed using a total of 282 billion tokens of biomedical text, including healthcare system-scale clinical text comprising 166 billion tokens from UF Health. We implemented a domain-adaptation method that integrates pre-training, model merging, a two-stage post-training approach, and agent-based reinforcement learning. An agentic retrieval-augmented generation (RAG) approach dynamically integrates time-sensitive clinical guidelines into the reasoning process. In a blind, randomized clinical evaluation conducted by five UF Health oncologists, GatorOnco significantly outperformed open-source LLMs (P < 0.01) and achieved expert-level performance comparable to UF Health oncologists. Compared with expert oncologists, GatorOnco received significantly higher ratings for readability (4.46 vs. 4.19, P < 0.01) and completeness (3.91 vs. 3.52, P < 0.01), while showing statistically comparable performance in correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999). These findings demonstrate that integrating agentic reasoning with large-scale domain adaptation can help bridge the gap for generative AI in high-stakes cancer treatment planning.