用于乳腺癌治疗建议的智能体系统
Agentic systems for breast cancer treatment recommendations
浏览论文内容
中文总结 AI 辅助
研究评估用于乳腺癌治疗建议的智能体LLM系统,用72个真实临床病例和1147个特定病例量表,比较七种流程,最佳配置全局得分为0.594±0.025,工具使用和智能体自主性影响各异,虽能生成相关建议,但用于无监督临床使用仍不足。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地被用于临床决策支持,但其在复杂肿瘤治疗规划中的可靠性仍不明确。我们使用72个从I期到IV期的真实临床病例和通过不对称信息量表生成(AIRG)产生的1147个特定病例量表,评估了用于生成乳腺癌治疗建议的智能体LLM系统。比较了七种流程,包括单LLM基线、工具增强系统和具有事实核查及自主子智能体生成的多智能体架构。最佳配置Claude Opus 4.8与D&C+SA流程的全局得分为0.594±0.025。工具使用和智能体自主性增加有不同影响。性能因临床领域和疾病阶段而异,肿瘤学家主导的错误分析揭示了持续存在的临床相关失败。这表明智能体LLM系统可生成临床相关的乳腺癌建议,但用于无监督临床使用仍不足。
英文摘要
Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer treatment recommendation generation using 72 real clinical cases across stages I to IV and 1,147 case-specific rubrics generated through Asymmetric Information Rubric Generation (AIRG), in which the rubric generator had access to real clinical decisions unavailable to the evaluated models. Seven pipelines were compared, including single-LLM baselines, tool-augmented systems, and multi-agent architectures with fact checking and autonomous subagent spawning. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 $\pm$ 0.025. Tool use and increased agent autonomy had mixed effects, improving performance in some settings but degrading it in others. Performance varied by clinical domain and disease stage, and oncologist-led error analysis revealed persistent clinically relevant failures, including incorrect or missing recommendations, flawed justifications, citation errors, outdated claims, and overconfidence. These findings suggest that agentic LLM systems can generate clinically relevant breast cancer recommendations, but remain insufficient for unsupervised clinical use.
发表机构
- Spesia(斯佩西亚)
- Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院)
- Laboratory of Artificial Intelligence Applied to Bioinformatics, SEPT, Universidade Federal do Paraná (UFPR)(巴拉那联邦大学人工智能应用于生物信息学实验室,SEPT)
- Pontifícia Universidade Católica do Paraná (PUCPR)(巴拉那天主大学)
机构由 AI 辅助整理,请以论文原文为准。